arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PATS:用于智能体强化学习的策略感知训练框架

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning

Yipeng Shi, Zhipeng Ma, Yue Wang, Qitai Tan, Yang Li, Peng Chen, Zhengzhou Zhu

arXiv 2607.21419首次发表:更新:

AI 中文总结

研究针对长期语言模型智能体强化学习中弱策略问题,提出以策略为中心的训练范式Pats,将技能作为动态训练框架,通过转换展开组为证据卡、特定任务评估调整上下文等改进策略,在多个任务上取得良好效果。

AI 中文摘要

在长期的语言模型智能体强化学习中,弱策略常重复类似失败,产生无信息的展开轨迹并限制有效策略优化。现有以技能为中心的方法通过优化、过滤或内化可复用技能来改进探索,但仍以技能本身为中心,而非作为对不断演变策略的适应性训练支持。为此,我们提出以策略为中心的训练范式,将技能重新构建为动态训练框架。我们的框架Pats将最新策略的展开组转换为证据卡,并使用特定任务评估来调整后续展开中使用的上下文。具体指导帮助弱策略完成具有挑战性的任务。随着策略改进,冗余上下文被修订或删除,以减少对明确指导的依赖,同时保留有用的展开变化。使用标准RLVR通过环境奖励优化策略,在部署时丢弃训练框架。在ALFWorld和WebShop上,Pats比强大的基线提高了18.6%。在七个搜索增强的问答基准测试中,它保持竞争力,同时使用的提示令牌比基线少32.1%。

英文摘要

In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting effective policy optimization. Existing skill-centric methods improve exploration by optimizing, filtering, or internalizing reusable skills. However, they remain centered on the skills themselves rather than being designed as adaptive training-time support for the evolving policy. To address this, we propose a policy-centric training paradigm that reframes skills as a dynamic training scaffold. Our framework, PATS, converts rollout groups from the latest policy into evidence cards and uses task-specific evaluation to adjust the context used in subsequent rollouts. Concrete guidance helps weak policies to complete challenging tasks. As policy improves, redundant context is revised or removed to reduce reliance on explicit guidance while preserving useful rollout variation. The policy is optimized with environmental rewards using standard RLVR, and the training scaffold is discarded at deployment. Across ALFWorld, WebShop, and seven search-augmented QA benchmarks, PATS achieves performance competitive with SOTA baselines while using 25%-50% fewer tokens.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑