发表机构
Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对智能体强化学习的混淆问题,提出OCSD方法,通过对比两类回放视图推导观测残差,在高不确定步骤调制GRPO更新,在三类任务和多模型规模上优于基线。
AI 中文摘要
大型语言模型智能体通常通过强化学习进行训练,使用的是稀疏的轨迹级奖励,这类奖励对单个标记应如何更新提供的指导有限。在线策略自蒸馏(OPSD)通过在特权回放视图下对生成的标记重新评分,以获取密集的标记级监督,从而解决了这一问题。然而,我们发现了一个混淆问题:所得的支持信号可能同时反映了回放视图中包含的特权信息和回放支架导致的分数变化,使得难以将支持信号明确归因于该特权信息。当未来环境观测作为特权信息时,这个问题尤为明显,因为回放这些观测需要重建一个扩展的支架,而这本身会干扰标记分数。为解决这种混淆,我们提出了观测校准自蒸馏(OCSD),它对比了两个结构匹配的回放视图——完整视图和观测消融视图,二者仅在是否存在实际未来观测上存在差异,以此推导观测残差,从而抵消回放支架共享的分数变化。OCSD随后将该残差应用于高不确定步骤的标记级GRPO更新,同时保留轨迹级更新方向。在ALFWorld、WebShop和Search-QA上针对三个Qwen3模型规模开展的实验表明,OCSD始终优于强大的基线模型。诊断分析进一步证实,校准后的残差与局部环境反馈的对齐效果更好。我们的代码可在该https URL公开获取。
英文摘要
Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated. On-Policy Self-Distillation (OPSD) addresses this by re-scoring generated tokens under a privileged replay view to obtain dense, token-level supervision. However, we identify a confounding issue: the resulting support may reflect both the privileged information contained in the replay view and score shifts induced by the replay scaffold, making it difficult to attribute the support specifically to that information. This issue is especially pronounced when future environment observations serve as privileged information, since replaying them requires reconstructing an extended scaffold that itself perturbs token scores. To resolve this confounding, we propose Observation-Calibrated Self-Distillation (OCSD), which contrasts two structurally matched replay views, Full and Observation-Ablated, differing only in whether the actual future observation is present, to derive an observation residual that discounts score changes shared by the replay scaffold. OCSD then applies this residual to modulate token-level GRPO updates at high-uncertainty steps, while preserving the trajectory-level update direction. Experiments on ALFWorld, WebShop, and Search-QA across three Qwen3 model scales show that OCSD consistently outperforms strong baselines. Diagnostic analyses further confirm that the calibrated residual aligns better with local environment feedback. Our code is publicly available at https://github.com/yiy1x/OCSD.