arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniOPSD:统一结果与事后反馈的智能体强化学习

UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning

Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Xiaofeng Han, Zelong Zheng, Haoyu Wu, Tianyu Fu, Chenxu Zhao, Minghui Wu, Guannan He, Changwei Wang

arXiv 2609.34810首次发表:更新:

发表机构

University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; Peking University; Mininglamp Technology; Key Laboratory of Computing Power Network and Information Security, Ministry of Education; Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences); Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science(中国科学院大学; 中国科学院自动化研究所; 北京大学; 明略科技; 计算力网络与信息安全教育部重点实验室;山东省计算中心(齐鲁工业大学(山东省科学院)); 计算力互联网与服务计算重点实验室,山东省计算机科学基础研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

UniOPSD通过自适应局部信用仲裁统一结果与事后反馈,解决智能体强化学习中信用分配问题,在多个基准上显著提升性能。

AI 中文摘要

强化学习已成为训练语言模型智能体的有效方法,但稀疏且延迟的结果奖励在长交互序列中为信用分配提供的指导有限。近期关于在策略自蒸馏(OPSD)的工作通过评估策略在特权训练时上下文下的采样响应,提供了补充性监督。然而,我们的诊断显示,结果反馈与事后反馈之间的正平均一致性伴随着显著的局部不一致,这引发了如何在每个决策中分配两者影响力的疑问。我们提出了UniOPSD(统一在策略自蒸馏),通过自适应局部信用仲裁统一这些反馈源。UniOPSD在共享交互锚点处,从环境回报和成功同伴的事后反馈中构建可比较的信用估计。历史一致性决定全局混合水平,而当前信号可用性和相对精度调整每个源在个别决策中的影响力。保留了回合级结果贡献,并通过有界令牌调制细化融合的步骤信用以用于策略优化。使用Qwen2.5-3B-Instruct和Qwen2.5-7B-Instruct,UniOPSD在ALFWorld上分别达到82.8%和83.6%的成功率,在WebShop上分别达到75.0%和82.0%的成功率,在Search-QA上分别达到45.3%和49.8%的聚合准确率。在3B WebShop上,UniOPSD比SDAR提高了7.0个百分点。我们的代码可在该https URL获取。

英文摘要

Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's sampled responses under privileged training-time context. However, our diagnostics show that positive average agreement between outcome and hindsight feedback coexists with substantial local disagreement, raising the question of how to allocate influence between them at each decision. We introduce UniOPSD (Unified On-Policy Self-Distillation), which unifies these feedback sources through adaptive local credit arbitration. UniOPSD constructs comparable credit estimates from environmental returns and successful-peer hindsight at shared interaction anchors. Historical agreement determines the global mixing level, while current signal availability and relative precision adjust each source's influence at individual decisions. The episode-level outcome contribution is retained, and bounded token modulation refines the fused step credit for policy optimization. With Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct, UniOPSD achieves ALFWorld success rates of $82.8\%$ and $83.6\%$, WebShop success rates of $75.0\%$ and $82.0\%$, and Search-QA aggregate accuracies of $45.3\%$ and $49.8\%$, respectively. On 3B WebShop, UniOPSD improves over SDAR by $7.0$ percentage points. Our code is available at https://github.com/Zenghuang-Fu/Uniopsd

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑