arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23651cs.SEcs.AI

适得其反的反馈:为什么小型语言模型智能体重复它们刚刚看到的失败调用

Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail

Esmail Gumaan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现小型语言模型智能体的失败工具调用反馈会适得其反,调用表面形式是主要原因,替换失败描述或使其不可达可减少重复,明确指令或删除失败尝试无效。

中文摘要 AI 辅助

智能体框架会在对话记录中记录失败的工具调用及其错误消息,并要求模型继续,假设错误是纠正性信息。我们对此进行了衡量,将失败记录的纠正增益定义为重新发出刚刚失败的动作的对数概率变化,我们在两个环境(模拟工具调用和MBPP程序修复)中测试的所有指令调优模型(6个检查点,1.35亿至17亿参数,4个模型家族)的增益均为负。按动作长度归一化后,该效应约为每个动作标记-1.03 nat,即每个标记的几率降低2.8倍,且在90%-100%的单个项目上均存在,而非仅在平均值上。在固定候选集上,重复失败调用的概率从0.06上升至0.54,且在失败后,贪心解码会在19%的项目上逐标记重现该调用,而失败前为0%。将同一调用与失败消息、成功消息或中性确认配对的反事实实验分离出两种效应:失败调用的表面形式占了83%的损害,而标记其失败的语义贡献很小,且在不同环境中符号不一致。问题出在智能体框架,而非模型对错误消息的理解,这也预测了哪些补救措施有效。将逐字调用替换为运行时生成的失败描述可消除76%的反转,且无标记成本;而让解码器无法访问之前失败的字符串则作用于同一术语。两种看似合理的补救措施无效:明确的“不要重复”指令使测量量保持不变;而从干净上下文删除失败尝试以重试(上下文污染的标准解决方案)是我们测量的最差重复框架,因为它恢复了导致失败的上下文。该研究在CPU上端到端运行,所有人工制品均已发布。

英文摘要

Agent harnesses record a failed tool call and its error message in the transcript and ask the model to continue, on the assumption that the error is corrective information. We measure whether it is. Defining the corrective gain of a failure record as the change in log-probability of re-emitting the action that just failed, we find the gain is negative for every instruction-tuned model we tested (6 checkpoints, 135M-1.7B, 4 families) in two environments: simulated tool calling and MBPP program repair. Normalised by action length the effect is about -1.03 nats per action token, a factor of 2.8 in the odds of each token, and holds on 90%-100% of individual items, not only on average. Over a fixed candidate set the probability of repeating the failed call rises from 0.06 to 0.54, and greedy decoding reproduces it token for token on 19% of items after the failure versus 0% before. Counterfactuals pairing the same call with a failure message, a success message, or a neutral acknowledgement separate two effects: the failed call's surface form accounts for 83% of the damage, while the semantic contribution of marking it failed is small and inconsistent in sign across environments. The problem is in the harness, not the model's grasp of error messages, and that predicts which remedies work. Replacing the verbatim call with a runtime-generated description of the failure removes 76% of the inversion at no token cost, and making previously-failed strings unreachable at the decoder acts on the same term. Two plausible remedies do not: an explicit "do not repeat" instruction leaves the measured quantity where it was, and deleting the failed attempt to retry from a clean context, the standard prescription for context contamination, is the worst harness we measured for repetition, because it restores the context that produced the failure. The study runs end to end on a CPU; all artefacts are released.

发表机构

  • Faculty of Computer Science and Mathematics(计算机科学与数学学院)
  • University of Passau(帕绍大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑