发表机构
University of Science and Technology of China; Xiaohongshu; The Hong Kong University of Science and Technology (HKUST-GZ)(中国科学技术大学; 小红书; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对在线策略蒸馏中分歧信号在长轨迹上失效的问题,提出GraphOPD,利用环境状态变化构建依赖图,以随机游走平稳分布评分步骤并融合分歧信号,在多个基准上显著提升智能体性能。
AI 中文摘要
在线策略蒸馏通过从教师策略提供密集的、步骤级的指导来对大型语言模型智能体进行后训练,适用于强化学习奖励稀疏且仅在每条轨迹结束时出现一次的场景。现有的实现方式根据每个步骤的师生分歧大小来分配这种指导,其依据是单轮直觉:较大的分歧标志着值得纠正的错误。当决策链跨越多个回合时,该规则会失效,因为早期的偏差会进入后续每个回合的上下文,使得教师策略与偏离的轨迹保持一致,而不是标记其成因,同时可互换的步骤会记录较大但结果无关的分歧。我们在一个智能体基准上证明了这一点,其中蒸馏最高分歧步骤相比随机选择没有带来一致的好处。为此,我们引入了GraphOPD,这是第一种将基于图的结构增强引入在线策略蒸馏以提升智能体能力的方法。它从环境自身的状态变化记录中读取哪些步骤启用了哪些后续步骤,不受污染师生差距的偏差影响,将其组织成依赖图,通过随机游走平稳分布对每个步骤进行评分,并将该结构信用与分歧信号融合成轨迹相对掩码,将监督集中在每条轨迹上最适合的步骤上。在ALFWorld、WebShop和SearchQA上,跨越三种模型规模和十一个基线,GraphOPD始终表现出竞争力,相比最强基线提升高达+5.8个百分点。执行回放审计进一步表明,该结构信用评分对真实因果影响的追踪远高于随机水平,两个融合信号各自独立必要,且同一信号可迁移到域外工具集成推理中。
英文摘要
On-policy distillation post-trains large language model agents by supplying dense, step-level guidance from a teacher policy when the reinforcement-learning reward is sparse and arrives only once per trajectory. Existing instantiations allocate this guidance by the size of the teacher-student divergence at each step, on the single-turn intuition that a large disagreement marks a mistake worth correcting. Once decisions chain over many turns, that rule misfires, since an early drift enters every later context both policies condition on, leaving the teacher consistent with the drifted trajectory instead of flagging its cause, while interchangeable steps register large but outcome-irrelevant divergences. We demonstrate this on an agentic benchmark, where distilling the highest-divergence steps brings no consistent benefit over random selection. To this end, we introduce GraphOPD, the first method to bring graph-based structural augmentation into on-policy distillation for agent capabilities. It reads which steps enabled which later ones from the environment's own record of state changes, immune to the drift that corrupts the teacher-student gap, organizes them into a dependency graph, scores each step by a random-walk stationary distribution over it, and fuses that structural credit with the divergence signal into a trajectory-relative mask concentrating supervision on each rollout's highest-aptitude steps. Across three model scales and eleven baselines on ALFWorld, WebShop, and SearchQA, GraphOPD shows competitive performance throughout, improving over the strongest baseline by up to +5.8 pp. An executed-replay audit further shows that this structural credit score tracks true causal impact far above chance, that both fused signals are independently necessary, and that the same signal transfers to out-of-domain tool-integrated reasoning.