arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07371cs.LGcs.CL

面向智能体强化学习的轨迹相对事后蒸馏

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning

Haoyu Zheng, Yun Zhu, Qing Wang, Wenqiao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对智能体强化学习中事后信号分配不清的问题,提出TRIAL框架,通过轨迹相对步分配优化监督信号,在WebShop、ALFWorld等任务上优于GRPO及多数基线方法,提升了任务性能。

中文摘要 AI 辅助

近期的智能体强化学习方法利用事后(hindsight)机制来补充稀疏的结果奖励。然而,一个完整的回合(rollout)可以产生大量此类信号,导致如何在各个决策步(turn)间合理分配这些信号尚不明确。我们提出了TRIAL,这是一个带有统一步对齐评分协议的轨迹相对事后蒸馏框架。对于每个决策步,TRIAL提取该决策所实现结果的结果视图,并在普通上下文和事后条件上下文下评估相同的响应。带符号的对数概率差值决定了 token 级监督的方向和局部强度,而步级幅度则在整个实现轨迹上进行联合归一化。所得的分配乘数的合格 token 加权平均值为1,在保持平均乘数不变的同时,将密集监督重新分配到各个决策步。在WebShop和ALFWorld上使用不同主干(backbone)的实验表明,TRIAL在主干、环境和评估指标的全部8种组合中均优于GRPO,且在6种组合中的6种上达到了6种方法中的最佳或并列最佳性能。在使用Qwen3-1.7B的WebShop上,TRIAL将成功率从56.4%提升至75.2%,任务得分从78.7%提升至85.7%。受控消融实验进一步表明,轨迹相对步分配相比单纯的密集事后蒸馏能带来显著增益。

英文摘要

Recent agentic reinforcement learning methods use hindsight to complement sparse outcome rewards. However, a completed rollout can yield many such signals, leaving their appropriate allocation across turns unclear. We introduce TRIAL, a trajectory-relative hindsight distillation framework with a unified turn-aligned scoring protocol. For each decision turn, TRIAL extracts an outcome view of that decision's realized consequence and evaluates the same response under ordinary and hindsight-conditioned contexts. The signed log-probability gap determines the direction and local strength of token-level supervision, while turn-level magnitudes are normalized jointly over the realized trajectory. The resulting allocation multipliers have an eligible-token-weighted mean of one, redistributing dense supervision across turns while fixing its average multiplier. Experiments on WebShop and ALFWorld with different backbones show that TRIAL outperforms GRPO across all eight combinations of backbone, environment, and evaluation metric, while achieving the best or tied-best performance among six methods on six of them. On WebShop with Qwen3-1.7B, TRIAL improves the success rate from 56.4% to 75.2% and the task score from 78.7% to 85.7%. Controlled ablations further show that trajectory-relative turn allocation provides substantial gains beyond those of dense hindsight distillation alone.

发表机构

  • Zhejiang University(浙江大学)
  • Shanghai AI Laboratory(上海人工智能实验室)
  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

↑