arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24114cs.AI

AHEAD:用于智能体强化学习的环境增强蒸馏自适应回溯

AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic RL

发表机构亚马逊云科技人工智能实验室 · 普渡大学
查看机构详情
  • AWS AI Labs(亚马逊云科技人工智能实验室)
  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiaolong Jin, Dingmin Wang, Vijay Lingam, Varun Kumar

首次发表
浏览论文内容

中文总结 AI 辅助

AHEAD是一种步骤感知的智能体强化学习框架,通过匹配不同监督源适配步骤类型,在多任务及多模型规模下提升了GRPO算法的任务成功率,减少了训练步骤并适配更严格的交互预算。

中文摘要 AI 辅助

用强化学习训练多轮大语言模型(LLM)智能体通常依赖轨迹级奖励,该奖励为每一步分配统一优势,无法识别哪些决策导致成功或失败。自蒸馏方法可通过为强化学习(RL)补充特权信息提供更细粒度的监督,但现有方法通常对每一步应用相同类型的特权信息,忽略了关键不对称性:常规步骤几乎不需要额外指导,而关键错误步骤则需要环境反馈无法单独提供的纠正方向。我们提出AHEAD,这是一种步骤感知框架,可将不同监督源与不同步骤类型匹配:教师接收所有步骤的环境反馈作为基于接地的密集信号,还接收LLM生成的错误步骤纠正提示,以补充环境反馈缺失的方向。该方法对标准GRPO算法的改动极小。在ALFWorld、WebShop和基于搜索的问答任务中,针对三种模型规模,AHEAD在7B参数下较GRPO提升了任务成功率(ALFWorld提升13.3个百分点,WebShop提升11.0个百分点),能在更少训练步骤中达到给定成功率,且比仅基于结果的RL和现有自蒸馏基线更能在更严格的交互预算内解决任务。

英文摘要

Training multi-turn LLM agents with reinforcement learning typically relies on trajectory-level rewards, which assign a uniform advantage to every step and cannot identify which decisions led to success or failure. Self-distillation methods can provide finer-grained supervision by augmenting RL with privileged information. However, existing approaches usually apply the same type of privileged information to every step in an indistinguishable manner, ignoring a key asymmetry: routine steps need little additional guidance, while critical error steps require corrective direction that environment feedback alone cannot provide. We propose AHEAD, a step-aware framework that matches different supervision sources to different step types. The teacher receives environment feedback on all steps as a grounded dense signal, and additionally receives LLM-generated corrective hints on error steps to supply the direction that environment feedback lacks. The method introduces minimal changes to the standard GRPO algorithm. Across ALFWorld, WebShop, and Search-based QA, and across three model scales, AHEAD raises task success (+13.3 points on ALFWorld and +11.0 on WebShop at 7B over GRPO), reaches a given success rate in fewer training steps, and solves tasks within tighter interaction budgets than outcome-only RL and prior self-distillation baselines.

补充信息

↑