arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

协调过程监督与智能体策略优化中基于结果的信用分配

Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

Jingxiao Yang, Wangjie Gan, Yingxuan Zhuang, Wenqi Zhang, Jintao Chen, Xuhong Zhang

arXiv 2608.31077首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体策略优化中监督与信用分配的差距,提出 TASPO 方法,通过聚合特权信息的动作级似然变化分配结果信用,在三个基准上较 GRPO 提升 10.6% 且泛化性更好。

AI 中文摘要

基于结果的强化学习为语言模型智能体提供可验证的反馈,但会将轨迹级优势均匀分配给所有决策,在长 horizon 交互中产生粗糙的信用分配。在线自蒸馏通过使用仅在训练期间可用的特权信息(PI)重新评估采样行为,提供更精细的监督。然而,精细的监督不一定等同于精细的信用分配:PI 诱导的似然变化描述了额外信息如何改变策略偏好,但不能直接确定可执行动作应如何继承已验证的任务结果,这造成了监督-信用差距。特权信号可能与当前交互状态无关、在与可执行动作不对齐的 token 粒度上操作,且缺乏强化所需的结果语义。我们提出 TASPO,它将特权监督转换为基于结果的动作信用分配。TASPO 从已验证的成功经验中构建适用于决策的 PI,在可执行动作级别聚合 PI 诱导的似然变化,并将相对动作支持转换为原始轨迹优势上的正、有界、均值保持权重。因此,已验证的结果决定了更新方向和平均规模,而 PI 仅在动作间重新分配信用。在三个智能体基准测试中,TASPO 比 GRPO 提升了 10.6%,并对未见过的任务表现出更好的泛化能力。进一步分析表明,TASPO 减少了监督不匹配,且动作级分配稳定了策略优化过程。这些发现为研究界提供了另一个有趣的视角。

英文摘要

Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-level advantage uniformly to all decisions, yielding coarse credit over long-horizon interactions. On-policy self-distillation offers finer supervision by re-evaluating sampled behavior with privileged information (PI) available only during training. However, fine-grained supervision is not necessarily fine-grained credit: PI-induced likelihood changes describe how additional information alters policy preference, but do not directly determine how an executable action should inherit the verified task outcome. This creates a supervision-credit gap. Privileged signals may be irrelevant to the current interaction state, operate at a token granularity misaligned with executable decisions, and lack the outcome semantics required for reinforcement. We introduce TASPO, which converts privileged supervision into outcome-grounded action credit. TASPO constructs decision-applicable PI from verified successful experience, aggregates PI-induced likelihood shifts at the executable-action level, and converts relative action support into positive, bounded, mean-preserving weights on the original trajectory advantage. Thus, the verified outcome determines the update direction and average scale, while PI only redistributes credit across actions. Across three agentic benchmarks, TASPO improves over GRPO by 10.6\% and generalizes better to unseen tasks. Further analysis indicates that TASPO reduces supervision mismatch and that action-level assignment stabilizes the policy optimization process. These findings offer the community another interesting perspective.

CommentsWork in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑