当历史失效:在误导性多轮历史下评估与改进工具使用
When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories
AI总结:
该研究针对多轮历史误导导致的工具调用智能体决策问题,提出可靠状态策略迁移方法,构建配对基准bench,在Qwen3模型上实现了优于多种基线的工具使用准确率。
AI中文摘要:
工具调用智能体从累积的对话和工具轨迹中推断任务状态。然而,在持续交互中,历史轨迹在对当前请求失去权威性后,可能仍保持结构有效性和语义合理性。我们表明,此类历史会劫持模型已有的策略:在Qwen3-1.7B模型上,污染会翻转原始轨迹下32.1%的正确决策,并频繁诱导对损坏实体或界面约定的重复使用。我们引入bench,这是一个配对基准,包含同步的原始(Original)、污染(Polluted)和专家状态(Oracle State)视图,保留系统策略、当前工具、最新请求和正确的下一步操作。11种保留正确结果的干预措施可分离出决策状态、实体绑定和界面执行中的故障,涵盖完整调用和非调用决策。我们进一步提出ours方法,通过对学生生成的前缀进行软监督,将专家(Oracle)条件下的教师策略迁移到仅观察污染历史的学生。在Qwen3-1.7B上,ours方法实现了87.0%的平衡工具使用准确率,优于Gold-SFT(66.3%)、专家序列蒸馏(82.3%)和离线策略令牌蒸馏(85.0%)。该方法可一致扩展:8B规模的教师将同样紧凑的1.7B学生提升至91.9%,而8B学生达到93.0%。所得策略还可迁移至干净历史、未见过的函数、独立重新生成的评估上下文、外部工具使用基准以及噪声多跳问答。这些结果确立了历史可靠性是工具使用的一个独特瓶颈,并证明可靠状态策略迁移是一种有效且可扩展的解决方案。
英文摘要:
Tool-calling agents infer task state from accumulated dialogue and tool traces. In persistent interactions, however, historical traces may remain structurally valid and semantically plausible after they cease to be authoritative for the current request. We show that such history can hijack a policy the model already possesses: on Qwen3-1.7B, pollution flips 32.1% of decisions that are correct under the original trajectory and frequently induces reuse of corrupted entities or interface conventions. We introduce bench, a paired benchmark with synchronized Original, Polluted, and Oracle State views that preserve the system policy, current tools, latest request, and gold next action. Eleven gold-preserving interventions isolate failures in decision state, entity binding, and interface execution across complete calls and non-call decisions. We further propose ours, which transfers an Oracle-conditioned teacher policy to a student observing only polluted history through soft supervision on student-generated prefixes. On Qwen3-1.7B, ours achieves 87.0% Balanced Tool-Use Accuracy, outperforming Gold-SFT (66.3%), Oracle sequence distillation (82.3%), and off-policy token distillation (85.0%). The method scales consistently: an 8B teacher raises the same compact 1.7B student to 91.9%, while an 8B student reaches 93.0%. The resulting policies further transfer to clean histories, unseen functions, independently regenerated evaluation contexts, external tool-use benchmarks, and noisy multi-hop question answering. These results establish history reliability as a distinct tool-use bottleneck and demonstrate reliable-state policy transfer as an effective and scalable solution.