arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18909cs.CLcs.AI

超越结果:面向高效智能体基准测试的双视角关系学习

Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

  • Hunyuan Team, Tencent(腾讯混元团队)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Xinshuai Guo, Junjie Wu, Dolly Deng, Yinghui Li, Hai-Tao Zheng, Suncong Zheng, Maxm Pan

AI总结:

针对智能体基准评估成本高的问题,提出DualViewEval方法,联合利用结果与过程关系压缩基准,在五个基准上以20个任务实现24-40倍压缩并提升预测精度,同时揭示智能体能力差异。

AI中文摘要:

智能体基准测试的评估成本远高于传统的大语言模型基准测试。因此,基准压缩是一种自然的解决方案,然而现有方法主要对任务-模型最终得分分布中的冗余进行建模,而这在智能体评估中至关重要。为弥补这一局限,我们分析了大规模轨迹数据,识别出六个与最终智能体性能系统性相关的互补过程信号。为了从完整视角解耦智能体性能冗余,我们提出了DualViewEval,一种智能体基准压缩方法,该方法联合利用结果关系与过程关系来学习一个精确大小的最小集,并预测完整基准的得分。在五个智能体基准和五个代表性基线上,DualViewEval在所有数据集上均取得了最佳结果。仅使用20个任务,它在APEX-Agents和BFCL上实现了24倍至40倍的压缩,相较于最强竞争者将平均绝对误差(MAE)降低了14.5%至28.2%,同时在SWE-bench Verified上相对于EssenceBench将Kendall's τ提升了高达7.2%。所选的最小集还揭示了不同智能体之间的能力差异,为高效的智能体模型开发提供了紧凑且具有诊断性的反馈。

英文摘要:

Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods primarily model redundancy in task--model final-score distributions, which is important in agentic evaluation. To address this limitation, we analyze large-scale trajectories and identify six complementary process signals that are systematically associated with final agent performance. To disentangle agent performance redundancy from a complete perspective, we propose DualViewEval, an agent benchmark compression method that jointly exploits outcome and process relations to learn an exact-size miniset and predict the full-benchmark scores. Across five agent benchmarks and five representative baselines, DualViewEval achieves the best results in all datasets. With only 20 tasks, it achieves $24\times$--$40\times$ compression on APEX-Agents and BFCL, reducing mean absolute error (MAE) by $14.5\%$--$28.2\%$ over the strongest competitors while improving Kendall's $τ$ by up to $7.2\%$ relative to EssenceBench on SWE-bench Verified. The selected minisets further reveal capability differences among different agents, providing compact and diagnostic feedback for efficient agentic model development.

↑