arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

默认陷阱:重新思考工具使用型LLM智能体中的计划评估

The Default Trap: Rethinking Plan Evaluation in Tool-Using LLM Agents

Xueqi Li, Jingjie Ning, Yibo Kong

arXiv 2609.36829首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示工具使用型LLM智能体评估中的“默认陷阱”风险,通过大规模实验证明需联合报告优先级响应性、呈现依赖的默认行为及任务成功率与成本。

AI 中文摘要

执行者可能对所提供的计划中优先级的变化做出强烈反应,但当移除一个与默认对齐的完整计划时,其在相同信息选择概率上的变化却很小。我们将后者被误解为对替代优先级响应较弱的风险称为“默认陷阱”。我们比较了优先考虑不同信息目标的配对计划,并以一个共享的无计划参考作为基准。一个会计恒等式将这些不同的行为对比联系起来。在160个选定的零售、航空和AgentDojo任务上的3200个决策窗口中,切换优先级强烈地改变了两个模型的选择,而两个计划与默认的对比则有所不同。在额外的2160个窗口中,反转账户列表顺序将默认目标选择改变了63.3至98.3个百分点;优先级切换效应在任一顺序下仍保持96.7至100.0个百分点。一项单独的3240窗口组件研究发现,在单一优先级句子的控制下效果显著,而附加文本的效果因组别和方向而异。最后,1080个完整任务片段产生了相对于无计划的成功差异,范围从-19.4到+8.3个百分点。所有零售和航空的成功区间都包含零;AgentDojo的结果描述了四个固定的应用世界。这些发现支持对优先级响应性、依赖于呈现方式的默认行为以及任务成功率和成本进行联合报告。

英文摘要

An executor can respond strongly to a change in a supplied plan's priority while showing a small change in the same information-selection probability when a default-aligned whole plan is removed. We call the risk of interpreting the latter as weak responsiveness to alternative priorities the default trap. We compare paired plans that prioritize different information targets with a shared no-plan reference. An accounting identity relates these distinct behavioral contrasts. Across 3,200 decision windows on 160 selected Retail, Airline, and AgentDojo tasks, switching priorities strongly redirects two models' choices, while the two plan-versus-default contrasts differ. In 2,160 additional windows, reversing account-list order shifts default target selection by 63.3-98.3 percentage points; priority-switching effects remain 96.7-100.0 points in either order. A separate 3,240-window component study finds strong control under single priority sentences, with effects of additional text varying by group and direction. Finally, 1,080 full-task episodes yield observed success differences of -19.4 to +8.3 points relative to no plan. All Retail and Airline success intervals include zero; AgentDojo results describe four fixed application worlds. These findings support joint reporting of priority responsiveness, presentation-dependent defaults, and task success and cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑