ArenaFlow:从轨迹排序到开放式智能体强化学习的分层信用传播
ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL
浏览论文内容
中文总结 AI 辅助
ArenaFlow提出分层信用传播框架,利用锦标赛排序和结构化反思,在步骤与技能层面传播奖励,以提升开放式智能体强化学习的性能。
中文摘要 AI 辅助
强化学习在可验证领域显著提升了大型语言模型(LLM)智能体的性能,但在开放式智能体任务中仍难以应用,因为此类任务的解决方案多样且难以获得可靠的标量奖励。近期成对评估方法通过用相对偏好替代逐点评分,缓解了奖励判别崩溃的问题。然而,这些方法仍将丰富的比较反馈压缩为单一的轨迹级奖励,掩盖了关键中间步骤,并阻碍了成功行为被整合为可复用技能。我们提出ArenaFlow,一种面向开放式智能体强化学习的分层信用传播框架。ArenaFlow利用基于锦标赛的相对排序来推导轨迹级奖励信号。每次比较还配备结构化反思评估,揭示三类监督信息:关键成功步骤、可复用策略技能以及检索技能的使用归因。在步骤层面,ArenaFlow根据锦标赛生存深度将轨迹级优势传播到高置信度的关键步骤,从而实现对局部推理行为的更精准优化。在技能层面,ArenaFlow通过群体级使用归因估计技能效用,并通过效用感知的更新、剪枝和检索维护全局技能记忆。由此产生的高效用技能进一步作为未来探索的策略先验。大量实验验证了ArenaFlow在开放式智能体任务上的有效性。
英文摘要
Reinforcement learning has substantially improved large language model (LLM) agents in verifiable domains, but remains difficult to apply to open-ended agent tasks, where solutions are diverse and reliable scalar rewards are hard to obtain. Recent pairwise evaluation methods alleviate reward discrimination collapse by replacing pointwise scoring with relative preferences. However, they still compress rich comparative feedback into a single trajectory-level reward, obscuring decisive intermediate steps and preventing successful behaviors from being consolidated into reusable skills. We propose ArenaFlow, a hierarchical credit propagation framework for open-ended agent reinforcement learning. ArenaFlow leverages tournament-based relative ranking to derive trajectory-level reward signals. Each comparison is further equipped with structured reflective evaluation, which reveals three types of supervision: pivotal success steps, reusable strategy skills, and usage attribution of retrieved skills. At the step level, ArenaFlow propagates trajectory-level advantages to high-confidence pivotal steps according to tournament survival depth, enabling more targeted optimization of local reasoning behaviors. At the skill level, ArenaFlow estimates skill utility from group-level usage attribution and maintains a global skill memory through utility-aware updating, pruning, and retrieval. The resulting high-utility skills further serve as policy priors for future exploration. Extensive experiments validate ArenaFlow's effectiveness on open-ended agent tasks.
发表机构
- Alibaba Token Hub, Alibaba Group(阿里巴巴通义实验室,阿里巴巴集团)
- Amap, Alibaba Group(高德,阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。