arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基准的基准:对话智能体的基准评估

Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents

Noam Koren, Roy Bar-Haim, Abigail Goldsteen

arXiv 2608.06329首次发表:更新:

AI 中文总结

该研究针对对话智能体基准质量评估不足的问题,提出基于LLM评判员的无参考评估框架,可区分基准质量等级,适用于合成与人工整理的基准,验证了其有效性与实用性。

AI 中文摘要

面向任务的对话智能体采用人工整理或自动生成的基准进行评估,但基准质量却很少被评估。劣质基准可能包含不一致的任务、过于简单的场景或有限的策略覆盖范围,导致评估结果不可靠。我们提出一种无参考框架,利用LLM评判员评估基准的一致性、复杂性和策略覆盖范围,同时提供弱点的可操作诊断。我们通过证明其与独立人工标注的一致性、评估不同能力LLM生成的基准以及经历可控质量下降扰动的基准,验证了该框架。在不同领域和评判员模型下,所提出的指标始终能区分基准的质量等级。我们进一步证明该框架适用于人工整理的基准。我们的框架为评估合成和人工整理的对话智能体基准提供了一种实用方法。

英文摘要

Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or limited policy coverage, leading to unreliable evaluations. We introduce a reference-free framework that uses LLM judges to assess benchmark consistency, complexity, and policy coverage, while providing actionable diagnostics of weaknesses. We validate the framework by demonstrating agreement with independent human annotations and by evaluating benchmarks generated by LLMs of varying capabilities, as well as benchmarks subjected to controlled quality-degrading perturbations. Across domains and judge models, the proposed metrics consistently distinguish between benchmark quality levels. We further demonstrate the framework's applicability to manually curated benchmarks. Our framework offers a practical approach for evaluating synthetic and manually curated conversational-agent benchmarks.

Comments15 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑