arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ARISE:适应智能体强化学习中不断演变的能力缺口

ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning

Kun Feng, Yuchen Fang, Yiyang Tan, Shuqi Gu, Yongxiang Zhao, Yu Liu, Xingyu Lu, Lintao Ma, Kan Ren

arXiv 2609.35532首次发表:更新:

AI 中文总结

ARISE提出自适应评分标准-技能协同进化框架,利用展开证据动态调整评估与训练,以应对智能体能力缺口演变,在SkillsBench和Terminal-Bench上提升任务性能与训练效率。

AI 中文摘要

随着一个长时程智能体通过经验不断改进,先前观察到的弱点可能会消退,而新的限制会浮现,持续改变其仍需学习的内容。然而,学习过程往往仍然与对这些需求的静态看法紧密相连:固定的行为标准和训练优先级可能变得与不断演变的智能体能力不一致,而稀疏的任务级反馈使得这种不一致更加难以察觉。即使识别出了能力缺口,当前策略的展开(rollout)也可能反复产生相同的失败,而不是探索更好的替代方案。为了解决这个问题,我们引入了自适应评分标准-技能协同进化(Adaptive Rubric-Skill Co-Evolution, ARISE),这是一个强化学习框架,利用展开证据来持续调整评估标准、探索指导和训练优先级。评分标准不断演变以奖励部分行为进展,而其配对的技能则被精炼并选择性激活,以引导探索朝向未解决的弱点。在这种协同进化之外,基于能力的自适应采样优先处理那些针对需要进一步改进行为的任务。在两个具有挑战性的长时程智能体基准测试SkillsBench和Terminal-Bench上的实验表明,ARISE成功提升了整体任务性能和训练效率。项目页面位于此https URL。

英文摘要

As a long-horizon agent improves through experience, previously observed weaknesses may recede while new limitations emerge, continually changing what it still needs to learn. Yet the learning process often remains tied to a static view of these needs: fixed behavioral criteria and training priorities can become misaligned with evolving agent capabilities, while sparse task-level feedback makes such misalignment more difficult to detect. Even when capability gaps are identified, rollouts from the current policy may repeatedly reproduce the same failures rather than explore better alternatives. To address this, we introduce Adaptive Rubric-Skill Co-Evolution (ARISE), a reinforcement learning framework that uses rollout evidence to continually adapt evaluation criteria, exploration guidance, and training priorities. Rubrics evolve to reward partial behavioral progress, while their paired skills are refined and selectively activated to guide exploration toward unresolved weaknesses. Alongside this co-evolution, capability-based adaptive sampling prioritizes tasks that target behaviors needing further improvement. Experiments on two challenging long-horizon agent benchmarks, SkillsBench and Terminal-Bench, demonstrate that ARISE successfully enhances both overall task performance and training efficiency. The project page is at https://foundation-model-research.github.io/ARISE .

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑