arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14400cs.CL

智能体评估中的策略漏洞:当策略模糊性伪装成智能体错误

Policy Loopholes in Agent Evaluation: When Policy Ambiguity Masquerades as Agent Error

  • Amadeus SAS(艾玛迪斯公司)

机构由 AI 辅助整理,请以论文原文为准。

Hongliu Cao

AI总结:

本研究揭示自然语言策略的模糊性可导致智能体评估分数不可靠,提出策略漏洞分类法,并强调策略规范质量决定评估质量上限,建议基准开发者在收集注释前审计策略。

AI中文摘要:

智能体基准测试评估策略合规性,但假设每条策略都决定了唯一正确的行动。自然语言策略可能通过沉默、歧义或矛盾违反这一假设,从而允许多种可辩护的解读,而单一的金标准轨迹无法捕捉这些解读。通过审计两个 $\ au^2$-bench 领域,我们开发了此类策略漏洞的分类法,并表明受影响的任务会产生不可靠的分数:它们以不同方式降低不同模型的分数,并使每个模型在重复试验中的一致性降低。跨领域比较揭示,可利用性既需要策略模糊性,也需要工具宽容性:当策略复杂性超过工具所能强制执行的范围时,智能体不一致地解决缺口,分数变得不可靠。策略规范质量设定了评估质量的上限。基准开发者应在收集金标准注释之前审计策略。

英文摘要:

Agent benchmarks evaluate policy compliance but assume each policy determines a unique correct action. Natural-language policies can violate this assumption through silence, ambiguity, or contradiction, admitting multiple defensible readings that a single gold trajectory cannot capture. Auditing two $τ^2$-bench domains, we develop a taxonomy of such policy loopholes and show that affected tasks produce unreliable scores: they lower scores across different models in different ways and make every model less consistent across repeated trials. A cross-domain comparison reveals that exploitability requires both policy ambiguity and tool permissiveness: when policy complexity exceeds what tools can enforce, agents resolve gaps inconsistently and scores become unreliable. Policy specification quality sets the ceiling on evaluation quality. Benchmark developers should audit policies before collecting gold annotations.

补充信息

↑