arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从任务结果训练LLM智能体的顾问

Training Advisors for LLM Agents from Task Outcomes

Sergei Polezhaev, Barys Liskavets, Ori Press, Alexander Golubev

arXiv 2610.09858首次发表:更新:

AI 中文总结

提出Caddie方法,通过强化学习从任务结果训练评论家为LLM智能体提供自然语言建议,显著提升多跳问答成功率并跨模型、跨域泛化。

AI 中文摘要

大型语言模型智能体通过将推理和工具调用与环境观察交错进行来处理多步任务。先前的工作表明,自然语言反馈可以帮助这些智能体在任务执行过程中修正其决策。我们引入了Caddie,一种训练评论家(critic)的方法,使其能够在智能体处理任务时提供自然语言分析和建议。与依赖步骤级标签或参考评论的方法不同,Caddie从智能体在收到评论家反馈后是否最终成功中学习。我们在保持基础模型冻结的同时,通过强化学习优化评论家。使用单一基础模型在多跳问答上进行训练,我们的Qwen3-4B评论家提高了四种不同规模和架构的基础模型(包括三种未用于评论家训练的基础模型)的成功率。在MuSiQue基准上,训练后的评论家将Qwen3-4B的成功率提高了超过25个百分点,超越了没有评论家的Kimi K3的性能。同一评论家还在域外交互式基准(包括τ³和DeepDive)上带来了收益,且无需额外训练。我们的结果表明,智能体可以在推理时决定何时向评论家寻求帮助,并且基于结果的评论家训练可以产生跨基础模型和任务域转移的指导。

英文摘要

Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution. We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task. Unlike approaches that rely on step-level labels or reference critiques, Caddie learns from whether the agent ultimately succeeds after receiving the critic's feedback. We optimize the critic through reinforcement learning while keeping the base model frozen. Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rates across four base models of different scales and architectures, including three not used during critic training. On the MuSiQue benchmark, the trained critic improves Qwen3-4B's success rate by more than 25 percentage points, surpassing the performance of Kimi K3 without a critic. The same critic also yields gains on out-of-domain interactive benchmarks, including $τ^3$ and DeepDive, with no additional training. Our results show that agents can decide when to seek help from a critic at inference time and that outcome-based critic training can produce guidance that transfers across base models and task domains.

Comments26 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑