基准测试的基准:评估对话智能体评测集
文章背景与核心概要
在任务导向型对话智能体的开发中,基准测试(Benchmarks)是衡量模型性能的核心工具。然而,目前学术界和工业界往往忽视了对这些基准测试本身质量的评估。如果基准测试存在任务逻辑不一致、场景过于简单或策略覆盖范围有限等问题,将直接导致模型评估结果的不可靠。
本文提出了一种无需参考答案(Reference-free)的评估框架,利用大语言模型(LLM)作为裁判,对基准测试的逻辑一致性、复杂度及策略覆盖度进行全面评估,并提供可操作的诊断建议。该研究通过与人类标注结果对比、评估不同能力LLM生成的基准,以及对基准进行受控质量扰动实验,验证了该框架的有效性。这一成果为评估合成及人工构建的对话智能体基准测试提供了一种实用且可靠的方法。
论文元数据 (Paper Metadata)
- arXiv ID: arXiv:2608.06329 [cs.CL]
- 学科分类: 计算与语言 (
cs.CL);人工智能 (cs.AI) - 提交日期: 2026年8月6日
- 篇幅: 15页
- 作者:
- Noam Koren
- Roy Bar-Haim
- Abigail Goldsteen
摘要 (Abstract)
任务导向型对话智能体通常使用人工策划或自动生成的基准进行评估,但基准本身的质量却很少受到评估。质量低劣的基准可能包含不一致的任务、过于简单的场景或有限的策略覆盖,从而导致评估结果不可靠。我们引入了一个无需参考答案的框架,利用大语言模型(LLM)裁判来评估基准的一致性、复杂度和策略覆盖度,同时提供针对弱点的可操作诊断。我们通过证明其与独立人类标注的一致性,以及评估由不同能力LLM生成的基准和经过受控质量降级扰动的基准,验证了该框架。在不同领域和裁判模型中,所提出的指标始终能够区分基准的质量水平。我们进一步证明了该框架对人工策划基准的适用性。我们的框架为评估合成及人工策划的对话智能体基准提供了一种实用的方法。
Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or limited policy coverage, leading to unreliable evaluations. We introduce a reference-free framework that uses LLM judges to assess benchmark consistency, complexity, and policy coverage, while providing actionable diagnostics of weaknesses. We validate the framework by demonstrating agreement with independent human annotations and by evaluating benchmarks generated by LLMs of varying capabilities, as well as benchmarks subjected to controlled quality-degrading perturbations. Across domains and judge models, the proposed metrics consistently distinguish between benchmark quality levels. We further demonstrate the framework's applicability to manually curated benchmarks. Our framework offers a practical approach for evaluating synthetic and manually curated conversational-agent benchmarks.
链接与资源 (Links & Resources)
- 全文访问:
- 查看 PDF
- TeX 源码
- DOI 链接
- 外部引用:
- Google Scholar
- Semantic Scholar
- NASA ADS