重试、切换还是弃权?通过受控错误注入学习策略感知的工具使用策略
Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection
浏览论文内容
中文总结 AI 辅助
该研究提出BENCH2ROBUST框架,结合BTM与课程强化学习,提升LLM智能体在工具故障下的稳健性,在零售任务中BTM可提稳健性16.8个百分点,二者结合达40.8-45.5%。
中文摘要 AI 辅助
使用工具的大语言模型(LLM)智能体通常在工具调用能可靠成功的环境中进行训练和评估,但部署的工具可能出现瞬时、持续或静默的故障。因此,稳健的恢复不仅需要重复重试:智能体可能需要重试同一路径、切换到替代路径,或识别出没有可行路径剩余。我们提出BENCH2ROBUST,这是一个将无故障工具使用基准转换为具有场景可控可解性的受控随机环境的框架,其中情节明确要求在可用路径耗尽后重试、切换或停止。我们使用BENCH2ROBUST研究两种互补干预措施:通过贝叶斯工具记忆(Bayesian Tool Memory,BTM)构建结构化运行时恢复上下文,以及课程控制的强化学习。在来自4个系列的7个模型和两个多轮基准系列中,工具故障产生了近乎普遍的稳健性差距。在保留的零售任务上,BTM在不重新训练的情况下将稳健性提高了多达16.8个百分点,而强化学习学习到互补的恢复行为,在没有推理时BTM的情况下仍然有益。结合两者在注入(故障)下达到40.8-45.5%的稳健性,同时保持无故障性能。这些结果表明,稳健的工具使用受益于结合特定环境的恢复知识与学习到的恢复行为。
英文摘要
Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed reliably, yet deployed tools can fail transiently, persistently, or silently. Robust recovery therefore requires more than repeated retries: an agent may need to retry the same path, switch to an alternative, or recognize that no viable path remains. We present BENCH2ROBUST, a framework that converts failure-free tool-use benchmarks into controlled stochastic environments with scenario-controlled solvability, where episodes explicitly require retrying, switching, or stopping after available paths are exhausted. We use BENCH2ROBUST to study two complementary interventions: structured runtime recovery context through Bayesian Tool Memory (BTM), and curriculum-controlled reinforcement learning. Across 7 models from 4 families and two multi-turn benchmark families, tool failures produce a near-universal robustness gap. On held-out Retail tasks, BTM improves robustness by up to 16.8 percentage points without retraining, while RL learns complementary recovery behavior that remains beneficial without inference-time BTM. Combining the two reaches 40.8-45.5% under injection while preserving failure-free performance. These results suggest that robust tool use benefits from combining environment-specific recovery knowledge with learned recovery behavior.
发表机构
- Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。