arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向基于技能的大语言模型智能体强化学习的双向上下文自蒸馏

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents

Tianjun Pan, Yuan Li, Hongda Wang, Linbo Jin, Mengfei Song, Lei Gao, Qiming Shi, Shaokang Fu, Jiarong Zhao, Chengyu Wang, Chengfu Huo

arXiv 2608.09555首次发表:更新:

AI 中文总结

本研究针对基于技能的LLM智能体技能利用不足问题,提出BCSD框架,通过双向上下文自蒸馏结合强化学习,在ALFWorld和WebShop上取得最优性能,验证了互补视角的有效性。

AI 中文摘要

外部自然语言技能为大语言模型(LLM)智能体提供了可复用、可编辑的复杂任务解决指导,但其有效性不仅取决于技能质量,还依赖于策略能否将提供的指导转化为合适的动作。然而,专门用于提升这种技能利用能力的方法仍未得到充分探索。实践中,基于技能的智能体通常以任务级奖励为核心的强化学习目标进行训练,这类目标提供的监督有限,难以捕捉策略利用所提供技能的效果差异。我们提出BCSD(双向上下文自蒸馏),一种结合自蒸馏与强化学习的框架,用于训练LLM智能体更有效地利用外部技能。与依赖单一特权上下文的现有自蒸馏方法不同,BCSD从两个互补的技能-上下文视角评估每个轨迹:增强视角引入更高级的元技能指导,简化视角则修剪通用指导以突出任务特定技能;二者互补的 token 级信号被结合起来以重新缩放强化学习优势。在ALFWorld和WebShop上的实验表明,BCSD在各模型规模下均实现了最强的整体性能,使智能体能更有效地利用外部技能; ablation研究进一步验证了增强和简化上下文视角的互补贡献,代码将被发布以确保完全可复现。

英文摘要

External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks. Yet their effectiveness depends not only on skill quality, but also on whether the policy can translate the provided guidance into appropriate actions. However, methods specifically designed to improve this skill-utilization ability remain largely underexplored. In practice, skill-based agents are commonly trained with reinforcement learning objectives centered on task-level rewards, which offer limited supervision and struggle to capture subtle differences in how effectively the policy uses the provided skills. We propose BCSD (Bidirectional Context Self-Distillation), a framework that combines self-distillation with reinforcement learning to train LLM agents to use external skills more effectively. Unlike prior self-distillation methods that rely on a single privileged context, BCSD evaluates each trajectory from two complementary skill-context views. The augmented view introduces higher-level Meta-Skill guidance, while the reduced view prunes general guidance to highlight task-specific skills. Their complementary token-level signals are combined to rescale the RL advantage. Experiments on ALFWorld and WebShop demonstrate that BCSD achieves the strongest overall performance across model scales, enabling agents to utilize external skills more effectively. Ablation studies further verify the complementary contributions of the augmented and reduced context views. Code will be released to ensure full reproducibility.

Comments9 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑