面向长周期工具使用智能体任务的高效强化学习
Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks
浏览论文内容
中文总结 AI 辅助
本文提出 SINKFLEX-RL 模块化训练系统,通过整合环境接口等技术,在 Tau2Bench 测试中提升验证奖励并降低注意力路径显存占用,实现长周期工具使用智能体的高效 RL 训练。
中文摘要 AI 辅助
长周期工具使用智能体需对用户目标、领域策略、工具调用、模拟器状态及延迟可验证奖励进行推理。强化学习(RL)适用于该场景,但多回合在线策略 rollout 会产生长上下文,而特定模型的注意力层可能需要自定义掩码和学习到的 sink 归一化。本文提出 SINKFLEX-RL,一种用于双控制工具使用环境中 RL 的模块化训练系统。该系统结合了兼容 Gymnasium 的环境包装器、VERL 风格的 rollout 数据流、无单独价值模型的组相对策略优化,以及 sink 感知的 FlexAttention 路径,该路径旨在在因果和滑动窗口掩码下保留特定模型的 sink 缩放。在初步的 Tau2Bench 零售运行中,验证奖励(mean@1)从训练初期的 0.25 上升到观测训练窗口后期的 0.44,同时训练分数和轨迹奖励代理也呈上升趋势。在固定配置的内存基准测试中,优化后的注意力路径在 4096 token 时将峰值 VRAM 从 28.06GB 降低至 22.52GB,降幅为 19.7%,并以 25.53GB 运行实测的 8192 token 配置,而 eager 基准则内存不足。这些结果表明,集成环境接口、RL 数据流和注意力内核设计对于实现内存可行的长周期智能体训练具有重要价值。
英文摘要
Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcement learning (RL) is a natural fit for this setting, but multi-turn on-policy rollouts create long contexts, while model-specific attention layers may require custom masks and learned sink normalization. We present SINKFLEX-RL, a modular training system for RL in dual-control tool-use environments. The system combines a Gymnasium-compatible environment wrapper, a VERL-style rollout dataflow, group-relative policy optimization without a separate value model, and a sink-aware FlexAttention path designed to preserve model-specific sink scaling under causal and sliding-window masks. In a preliminary Tau2Bench retail run, validation reward (mean@1) rises from 0.25 early in training to $0.44$ later in the observed training window, while training-score and trajectory-reward proxies also trend upward. In a fixed-configuration memory benchmark, the optimized attention path reduces peak VRAM from 28.06GB to 22.52GB at 4096 tokens, a $19.7\%$ reduction, and runs the measured 8192-token configuration using $25.53$~GB where the eager baseline runs out of memory. These results illustrate the value of integrating environment interfaces, RL dataflow, and attention-kernel design for memory-feasible long-horizon agent training.
发表机构
- Capital One(第一资本金融公司)
- AI Foundations(人工智能基础部门)
机构由 AI 辅助整理,请以论文原文为准。