arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DelegationBench:衡量AI代理何时应在行动前询问

DelegationBench: Measuring When AI Agents Should Ask Before Acting

Shiva Pochampally

arXiv 2610.05532首次发表:更新:

AI 中文总结

DelegationBench通过156个场景和匹配对测试AI代理行动前询问的评分可信度,发现一致性分数存在误导,并建议分别报告属性而非单一分数。

AI 中文摘要

能够发送电子邮件、编辑文件和进行购买的AI代理必须决定何时自主行动以及何时先与用户确认。这种决策通常通过向模型展示一个提议的行动,询问其是否应继续,并根据与人工标签的一致性进行评分来评估。我们引入DelegationBench来测试此类分数是否可信。它包含156个场景,每个场景有四种可能的响应(执行、请求许可、请求缺失信息、拒绝),且大多数场景以匹配对的形式出现,这些匹配对仅改变一个特征:行动是否被请求、风险大小、是否可撤销或谁会看到它。在来自五个模型家族的十个模型中,一致性分数在三个方面产生误导。一个简单的关键词规则(我们在看到基准后编写)与我们的标注者的一致性高于八个模型,但其决策在48个匹配对中仅变化了9个。以等价方式提出同一问题会使模型执行行动的频率变化高达52.5个百分点。并且,每个模型在使用工具执行任务时比仅评判提议行动时更少停下来询问用户。当规则被明确陈述时,相同的模型几乎完美地遵循它们,因此这些差距并非由普遍无法遵循规则来解释。我们发布了基准和评估工具,并建议分别报告这些属性,而不是作为一个单一分数。

英文摘要

AI agents that send emails, edit files, and make purchases must decide when to act on their own and when to check with the user first. This decision is usually evaluated by showing a model a proposed action, asking whether it should proceed, and scoring agreement with human labels. We introduce DelegationBench to test whether such scores can be trusted. It has 156 scenarios with four possible responses (act, ask for permission, ask for missing information, refuse), and most scenarios come in matched pairs that change a single feature: whether the action was requested, what is at stake, whether it can be undone, or who will see it. Across ten models from five families, agreement scores mislead in three ways. A simple keyword rule, which we wrote after seeing the benchmark, agrees with our annotators more often than eight of the models, yet its decision changes in only 9 of 48 matched pairs. Equivalent ways of asking the same question change how often a model acts by up to 52.5 percentage points. And every model stops to ask the user less often when it must carry out the task with tools than when it judges a proposed action. When rules are stated explicitly, the same models follow them almost perfectly, so the gaps are not explained by a general inability to follow rules. We release the benchmark and evaluation tools and recommend reporting these properties separately rather than as one score.

Comments35 pages, 7 figures. Code and data: https://github.com/PieLord757/delegation-bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑