arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RewardWeaver:通过自进化奖励适应实现语言智能体的长时程交互学习

RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation

Hengbo Xiao, Boyao Zhang, Purui Liu, Yuxuan Zheng, Haoran Yin, Haibo Liu, Fan Zhang

arXiv 2610.10120首次发表:更新:

发表机构

University of Science and Technology of China; Peking University; Tianjin University(中国科学技术大学; 北京大学; 天津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出自进化奖励适应框架RewardWeaver,通过动态调整过程奖励应对长时程交互中的稀疏反馈,在多个基准上取得SOTA结果。

AI 中文摘要

基于可验证奖励的强化学习(RLVR)在任务结果能够可靠评估的领域取得了显著进展,但长时程交互仍然具有挑战性,原因在于稀疏的终端反馈和困难的信用分配。过程奖励提供了更密集的监督,然而与训练最相关的能力可能会随着策略的演化而改变:一个容易评估或频繁欠缺的行为未必是当前限制任务成功的瓶颈。我们提出了RewardWeaver,一个面向长时程交互中语言智能体的自进化奖励适应框架。RewardWeaver维护一个经过验证的能力空间,其中已接纳的Rubrics(评估标准)的语义保持固定,并闭环连接策略优化、任务评估、失败归因和奖励适应。在每个训练阶段之后,它对低结果轨迹执行基于结果的后向归因,聚合反复出现且受策略控制的能力瓶颈,并动态选择相应的过程奖励用于下一阶段。现有能力空间未覆盖的反复失败会触发一个独立的、受控的扩展过程。我们在SOTOPIA、Amazon?HistoryPrice以及新构建的Sales Benchmark上评估了REWARDWEAVER。在社交互动、双边谈判和特定领域销售中,REWARDWEAVER均取得了新的最先进(SOTA)结果。消融实验进一步证明了动态奖励分配、基于失败的归因以及已接纳能力的稳定语义的重要性。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in domains where task outcomes can be reliably evaluated, but long-horizon interaction remains challenging due to sparse terminal feedback and difficult credit assignment. Process rewards provide denser supervision, yet the capabilities most relevant for training can change as the policy evolves: a behavior that is easy to evaluate or frequently deficient need not be the bottleneck currently limiting task success. We introduce RewardWeaver, a self-evolving reward adaptation framework for language agents in long-horizon interaction. RewardWeaver maintains a validated capability space in which the semantics of admitted Rubrics remain fixed, and closes the loop between policy optimization, task evaluation, failure attribution, and reward adaptation. After each training stage, it performs outcome-grounded backward attribution on low-outcome trajectories, aggregates recurrent and policy-controlled capability bottlenecks, and dynamically selects the corresponding process rewards for the next stage. Recurrent failures not covered by the existing capability space trigger a separate, controlled expansion procedure. We evaluate REWARDWEAVER on SOTOPIA, Amazon?HistoryPrice, and a newly constructed Sales Benchmark. Across social interaction, bilateral bargaining, and domain-specific sales, REWARDWEAVER establishes new state-of-the-art (SOTA) results. Ablations further demonstrate the importance of dynamic reward allocation, failure-grounded attribution, and stable semantics for admitted capabilities.

Comments23 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑