arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体是系统,而非模型:重新思考智能体评估

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata

arXiv 2610.01618首次发表:更新:

发表机构

Technical University of Munich; Helmholtz Munich; Inria; École normale supérieure; CNRS; PSL Research University(慕尼黑工业大学; 亥姆霍兹慕尼黑; 法国国家信息与自动化研究所; 巴黎高等师范学院; 法国国家科学研究中心; 巴黎文理研究大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出智能体应被视为可配置系统而非固定模型,通过四项科学任务基准研究配置要素影响,发现信息提供影响最大,并建议通过系统设计而非提示实现期望行为。

AI 中文摘要

智能体评估日益超越单一成功率,报告成本、一致性和鲁棒性等指标。然而,这些评估通常将智能体本身视为固定不变的。在实践中,智能体是一个可配置的系统:用户决定提供什么信息、允许其运行多长时间以及使用哪个模型,而这些选择中的每一个都可能改变其表现的优劣和一致性。我们在一个包含四项科学任务的新基准上研究了这些选择,其中编码智能体必须找到并正确操作一个已发布的专业模型。我们调查了智能体配置的五个部分:任务信息、推理、自我验证、时间预算和主干模型。我们发现存在显著的运行间变异性,约54%的结果方差来自重复相同配置而非更改配置。在不同配置中,提供给智能体的信息影响最大,超过了时间预算和模型规模,同时还能降低成本并改善校准。配置选择之间也存在交互:额外的时间仅在智能体拥有足够信息或足够强大的模型来利用它时才有帮助。最后,基于轨迹的智能体行为分类法揭示,提示智能体验证其答案对其验证行为影响甚微,而提供专门的验证工具则显著改变了该行为。这些结果表明,智能体应作为可配置系统本身进行评估,且某些期望行为通过系统实现比通过提示请求更有效。我们发布了该基准和超过18,000条智能体轨迹。

英文摘要

Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific tasks, where a coding agent must find and correctly operate a published specialist model. We investigate five parts of the agent's configuration: task information, reasoning, self-verification, time budget, and backbone model. We find substantial run-to-run variability, with approximately 54% of the outcome variance coming from repeating the same configuration rather than changing it. Across configurations, the information provided to the agent has the largest effect, exceeding both time budget and model size, while also reducing cost and improving calibration. Configuration choices also interact: additional time helps only when the agent has sufficient information or a capable enough model to use it. Finally, a trajectory-based taxonomy of agent behavior reveals that prompting an agent to verify its answer has little effect on its verification behavior, whereas providing a dedicated verification tool changes that behavior substantially. These results suggest that agents should be evaluated as configurable systems themselves, and that some desired behaviors are more effectively implemented in the system than requested through prompting. We release the benchmark and more than 18,000 agent trajectories.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑