arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33391cs.LGcs.AI

超越时间戳:面向长时程智能体的决策对齐在线策略蒸馏

Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents

Mingju Chen, Can Lv, Jinrong Liu, Huan Zhang, Heng Chang, Shiji Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

针对长时程智能体强化学习中特权监督与决策时间戳错位的问题,提出AlignOPSD方法,通过决策对齐监督修正和半马尔可夫层次信用分配,在多个基准上显著优于现有方法。

中文摘要 AI 辅助

具有可验证奖励的强化学习(RLVR)通常依赖于稀疏的结果奖励,为长时程智能体提供粗略的监督。在线策略自蒸馏(OPSD)通过密集的特权反馈补充了这一信号。然而,我们发现了“决策-时间戳不匹配”问题:特权指导可能与学生的功能决策不对齐,因为相应的决策可能发生在不同的时间步,而学生自身的决策可能跨越多个时间步,而非绑定于单个时间戳。因此,时间戳局部的监督可能在上下文和信用分配的时间范围上均产生错位。为解决这一不匹配,我们引入了AlignOPSD,遵循“先对齐监督,再分配信用”的原则。决策对齐的监督修正(Decision-Aligned Supervision Rectification)在功能匹配的上下文中,对同一学生采样的响应在兄弟轨迹(sibling rollouts)间重新评分,以校准局部教师证据。半马尔可夫层次信用分配(Semi-Markov Hierarchical Credit Assignment)则从对应关系变化中推导出可变时长的决策跨度,并使用修正后的证据在跨度及其组成轮次间分配基于结果的信用。我们使用Qwen2.5-3B和Qwen2.5-7B在ALFWorld、WebShop和Search-QA上评估AlignOPSD,并与代表性基线进行比较。AlignOPSD在所有八个骨干-聚合指标比较中均优于GRPO和StepOPSD,相比GRPO提升了5.5%至8.7%,并在六项中排名第一。额外的分析考察了两个对齐阶段以及任务间的超参数敏感性。我们的代码可在以下网址获取:此https URL。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify \emph{Decision--Timestamp Mismatch}: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce \textsc{AlignOPSD}, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate \textsc{AlignOPSD} with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. \textsc{AlignOPSD} outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 \% and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at https://github.com/mingju-c/Align-OPSD

发表机构

  • Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(北京未来区块链与隐私计算高精尖创新中心)
  • Tsinghua University(清华大学)
  • Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑