arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多步工具增强智能体中的绑定漂移

Binding Drift in Multi-Step Tool-Augmented Agents

Rahul Suresh Babu, Shashank Indukuri

arXiv 2607.18316首次发表:更新:

AI 中文总结

研究多步工具增强智能体中实体绑定随时间的变化,区分绑定漂移与错误传播。通过实验发现实体锁定会放大错误操作,实用的重新验证器可减少错误,自然设置下基线智能体有漂移现象,且持久性和重新验证不可互换。

AI 中文摘要

工具增强语言模型智能体在外部系统上执行多步工作流程,先解析一个实体,然后在后续步骤中对其进行操作。先前工作表明,在单步操作中,智能体有24%-26%的时间会将正确工具绑定到错误实体。本文研究实体绑定随时间的变化情况:它们是保持正确、悄悄漂移到不同实体,还是从一开始就错误并传播和加剧?我们将绑定漂移(第一步正确,之后错误)与错误传播(第一步错误,延续下去)区分开来,并在不相交的工作流程集上对它们进行评分,以免两者混淆。在一个受控的多步测试平台(200个工作流程、580个有实体绑定评分的步骤、四个企业领域、八个从小型到前沿的模型后端)中,我们发现:(1)在受控错误注入下,实体锁定(直观的“保留第一个绑定”修复方法)将错误操作从907次放大到2746次(3.0倍;自展95%置信区间[2.8, 3.3]),因为它将植入的错误实体忠实地带入每个后续步骤;(2)在受影响最大的模型(Claude Opus 4.5)上,放大倍数达到8.5倍;(3)一个基于实用语言模型的重新验证器(单次低成本的第二个模型调用,重新读取原始指令)将错误操作减少了79%(0.21倍;置信区间[0.18, 0.25]),将差距缩小到与神谕上限(0.20倍)相差不到1个百分点;(4)在自然(未注入)设置中,基线智能体在18%的合格工作流程上漂移,每步错误率随步骤增加。持久性和重新验证不可互换:消除漂移的防御措施可能会使传播恶化,而实用的重新验证器几乎与神谕恢复效果相当。

英文摘要

Tool-augmented language-model agents execute multi-step workflows over external systems, resolving an entity once and then acting on it across subsequent steps. Prior work shows that in single-step actions, agents select the correct tool but bind it to the wrong entity 24-26% of the time. We study what happens to entity bindings over time: do they stay correct, silently drift to a different entity, or, if wrong from the start, propagate and compound? We formalize binding drift (correct at step 1, wrong later) as distinct from error propagation (wrong at step 1, carried forward), and score them on disjoint workflow sets so the two cannot be conflated. In a controlled multi-step testbed (200 workflows, 580 entity-binding-scored steps, four enterprise domains, eight model backends spanning small to frontier), we find: (1) under controlled error injection, an entity lock (the intuitive "persist the first binding" fix) amplifies wrong actions from 907 to 2,746 (3.0x; bootstrap 95% CI [2.8, 3.3]), because it faithfully carries the seeded wrong entity into every later step; (2) the amplification reaches 8.5x on the most affected model (Claude Opus 4.5); (3) a practical LLM-based re-verifier (a single cheap second model call re-reading the original instruction) reduces wrong actions by 79% (0.21x; CI [0.18, 0.25]), closing the gap to within 1 percentage point of an oracle upper-bound (0.20x); and (4) in the natural (non-injected) setting, baseline agents drift on 18% of eligible workflows, with the per-step error rate rising across steps. Persistence and re-verification are not interchangeable: a defense that eliminates drift can worsen propagation, and a practical re-verifier nearly matches oracle recovery.

Comments14 pages, 5 tables, 1 figure. Equal contribution by both authors. Code and data: https://github.com/shashank-indukuri/binding-drift

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑