arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08077cs.AIcs.CLcs.IRcs.LG

自我反思蒸馏:将事后经验转化为先验预见

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

Haoxiang Zhang, Qinglin Chen, Hiroaki Hayashi, Zhuofeng Li, Siming Zhang, Jiaxin Zhang, Jixuan Chen, Fang Wu, Pan Lu, Silvio Savarese, Julian McAuley, Chien-Sheng Wu

首次发表
浏览论文内容

中文总结 AI 辅助

提出自我反思蒸馏(SRD),利用事后轨迹知识训练交互前预见,补充RLVR,在奖励均匀场景下显著提升成功率,最高达24.2个百分点。

中文摘要 AI 辅助

基于可验证奖励的强化学习(RLVR)主要通过交互后的标量结果奖励将智能体经验转化为学习信号。然而,对于群体相对目标,当所有轨迹获得相同奖励时,该信号会消失,尽管这些轨迹可能揭示关于任务要求和智能体失败方式的有用信息。我们提出一个补充性问题:事后经验能否教会智能体在行动前本可预见的内容?我们引入前瞻性学习(prospective learning),利用事后经验来监督交互前视角的预见预测,并通过自我反思蒸馏(SRD)实例化该方法。直觉上,一条完整的轨迹揭示了原本有用的知识和应避免的陷阱;SRD 将这种特权事后知识蒸馏为同一策略的轨迹盲预见。预见仅作为训练目标,推理时无需显式生成。在10个工具集成推理和长程智能体任务中,SRD 补充 RLVR 和自蒸馏基线,带来高达24.2个百分点的提升。其优势在奖励对比稀缺时尤为显著:当37%至98%的轨迹组在不同模型规模下奖励均匀时,SRD 仍能从采样轨迹中利用学习信号。在2B设置中,98%的组为全失败,RLVR 训练最终成功率为0.0%,而在相同轨迹预算下添加 SRD 达到60.6%。我们的结果表明,事后智能体经验不仅有助于评估或改进行为,还有助于在可用交互之前塑造预测性表征。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-interaction view, and instantiate it with Self-Retrospection Distillation (SRD). Intuitively, a completed trajectory reveals knowledge that would have been useful and pitfalls that should be avoided; SRD distills this privileged hindsight into trajectory-blind foresight of the same policy. Foresight serves only as a training target and need not be explicitly generated at inference time. Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD complements RLVR and self-distillation baselines with gains of up to 24.2 pp. Its advantage is especially pronounced when reward contrast is scarce: when 37--98% of rollout groups are reward-uniform across model scales, yet SRD can still exploit learning signal from sampled trajectories. In the 2B setting, where 98% of groups are all-failure, the RLVR training ends up at 0.0% success, while adding SRD reaches 60.6% under the same rollout budget. Our results suggest that post-hoc agent experience is useful not only for evaluating or improving behavior, but also for shaping predictive representations before available interaction.

发表机构

  • Salesforce AI Research(Salesforce人工智能研究院)
  • UC San Diego(加州大学圣地亚哥分校)
  • Texas A&M University(得克萨斯农工大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

↑