arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14945cs.AIcs.CL

信任并不足够:智能体强化学习中策略内自蒸馏的影响校准

Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL

Qizhen Lan, Xi Xiao, Xiangchen Guan, Mengchen Fan, Moule Lin, Jung Im Choi, Lijing Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

针对策略内自蒸馏的信任-效用不匹配问题,提出ICSD方法,在多基准任务上提升语言智能体性能,降低反向令牌的教师支持占比并提高梯度兼容性。

中文摘要 AI 辅助

策略内自蒸馏(OPSD)为语言智能体提供来自特权自教师的、基于其自身轨迹的密集令牌级监督。现有方法主要通过教师信任分配该监督,但信任无法揭示强调某令牌是否支持当前策略目标,我们将此称为信任-效用不匹配,并提出自蒸馏影响校准(ICSD)。对于每个受监督令牌,ICSD测量其重要性加权强化学习(RL)替代贡献对教师导向输出扰动的一阶响应。批量自适应校准将此非平稳信号转换为有界分配权重,同时保留每个动作回合内的原始辅助损失总量。这些分离的权重仅影响蒸馏损失,无需额外模型前向传播。在ALFWorld、WebShop和Search-QA上,ICSD在15亿至70亿参数的两个模型系列中,基于组相对策略优化(GRPO)和组内组策略优化(GiGPO),相较于仅信任分配,提升了所有匹配的聚合指标。在70亿参数规模下,其达到96.1%的ALFWorld成功率和93.1的WebShop得分。冻结批量分析显示,ICSD将分配给与目标相反令牌的教师支持质量从60.1%降至37.8%,并将与RL梯度的余弦兼容性提高0.192。配套代码仓库可在此httpsURL获取。

英文摘要

On-policy self-distillation (OPSD) gives language agents dense token-level supervision from a privileged self-teacher on the policy's own trajectories. Existing methods allocate this supervision mainly by teacher trust, but trust does not reveal whether emphasizing a token supports the current policy objective. We call this the trust-utility mismatch and introduce Influence Calibration for Self-Distillation (ICSD). For each supervised token, ICSD measures the first-order response of its importance-weighted RL surrogate contribution to a teacher-directed output perturbation. Batch-adaptive calibration converts this non-stationary signal into a bounded allocation weight while preserving the original auxiliary-loss mass within each action turn. These detached weights affect only the distillation loss and require no additional model pass. Across ALFWorld, WebShop, and Search-QA, ICSD improves all matched aggregate metrics over trust-only allocation under Group Relative Policy Optimization (GRPO) and Group-in-Group Policy Optimization (GiGPO), across two model families spanning 1.5B to 7B. At 7B, it reaches 96.1% ALFWorld success and a WebShop score of 93.1. Frozen-batch analyses show that ICSD reduces teacher-supported mass assigned to objective-opposed tokens from 60.1% to 37.8% and raises cosine compatibility with the RL gradient by 0.192. A companion repository is avail- able at https://github.com/lanqz7766/Influence-Calibration-for-On-Policy-Self-Distillation-in-Agentic-RL.

↑