arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体基准测试中的双重测量混杂:去脚手架、真值评分与超越均值的可靠性

The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean

Yonghong Zhang, Shadi Motaali, Vu Phong Dinh, Avin Piroutiniya, Jorge E. López de Vergara, Luis de Pedro, Ricardo Correia, Isabel M. Parra, Yong Xie

arXiv 2609.09218首次发表:更新:

AI 中文总结

针对智能体基准测试中的双重测量混杂问题,提出审计与修复协议,通过去脚手架、真值评分和可靠性指标,在ComtradeBench上揭示模型真实能力差异,提升评估有效性。

AI 中文摘要

智能体基准测试越来越多地用于比较大型语言模型(LLM)并指导部署决策,然而基准分数只有在衡量模型能力而非评估流程属性时才有意义。我们识别出一个双重测量混杂:执行关键决策由固定的脚手架而非模型完成,而评分器使用的评估标准可能无法反映任务正确性。我们在一个测量理论框架内统一了这些问题,该框架刻画了基准分数何时可被解释为模型能力的证据,并实例化了一个审计与修复协议,该协议(i)将执行关键决策从脚手架转移给模型,(ii)用基于种子的真值评分取代基于形状的评估,以及(iii)通过最坏情况和尾部风险指标报告超越均值的可靠性。在ComtradeBench上的实验表明,联合干预将几乎平坦的排行榜转变为可靠性谱系,该谱系区分了不同种子下的平均性能和鲁棒性。将审计应用于现有基准进一步表明,评分器有效性是基准特定的,而脚手架所有权在我们探测的任何地方都是一个不受控制的轴。我们的结果表明,基准分数应与其脚手架水平、评分标准和可靠性概况一起解读,为更有效的LLM智能体评估提供了一个实用框架。

英文摘要

Agent benchmarks are increasingly used to compare large language models (LLMs) and guide deployment decisions, yet benchmark scores are meaningful only if they measure model capability rather than properties of the evaluation pipeline. We identify a double measurement confound: execution-critical decisions are performed by a fixed scaffold instead of the model, while the scorer evaluates outputs using criteria that may not reflect task correctness. We unify these issues within a measurement-theoretic framework that characterizes when benchmark scores can be interpreted as evidence of model capability, and instantiate it with an audit-and-repair protocol that (i) transfers execution-critical decisions from the scaffold to the model, (ii) replaces shape-based evaluation with seeded ground-truth scoring, and (iii) reports reliability beyond the mean through worst-case and tail-risk metrics. Experiments on ComtradeBench show that the joint intervention transforms a nearly flat leaderboard into a reliability spectrum that distinguishes both average performance and robustness across seeds. Applying the audit to existing benchmarks further shows that scorer validity is benchmark-specific, whereas scaffold ownership is an uncontrolled axis wherever we probed it. Our results suggest that benchmark scores should be interpreted together with their scaffolding level, scoring criterion, and reliability profile, providing a practical framework for more valid evaluation of LLM agents.

Comments31 pages, 18 figures; supplementary material included as appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑