arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HARTS:基于任意展开树的混合注意力模型的高效智能体强化学习

HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees

Boyuan Meng, Peihua Bao, Hong Liu, Xiaowei Zhu, Chao Wang, Gen Li, Zhenxuan Pan

arXiv 2608.28158首次发表:更新:

发表机构

Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

HARTS是首个针对任意展开树前缀共享的混合注意力模型智能体RL系统,通过分块线性注意力算法等技术实现了4.81-4.87倍的训练加速,数值与奖励表现与基线相当。

AI 中文摘要

智能体强化学习(RL)常产生带有共享历史的不规则展开树,训练时从根到叶的轨迹会独立重新计算这些共享前缀。现有系统主要针对全注意力模型,缺乏与激活重计算兼容的密集、可微混合注意力执行。我们提出HARTS(混合注意力树结构RL),它在经过前缀压缩后的紧凑令牌工作中,联合规划微批次、数据并行(DP)副本分配和微批次槽调度。对于分块线性注意力,线性时间算法协调分块边界状态恢复与重放,在我们的打包执行模型下产生最少的顺序线性注意力调用次数。HARTS保留轨迹式训练的分块状态划分:不重复投影、MLP/MoE计算或最终输出,仅执行有界状态重放以实现数值对齐。每一轮,HARTS将所有分支打包为一个调用,通过可微状态交接传播梯度,支持激活重计算,并恢复每个令牌的对数概率。对于确定性、无令牌丢弃的top-k MoE路由,语义多重性恢复MoE目标的令牌权重和负载统计。现有RL目标保留其接口。据我们所知,HARTS是首个在真实混合注意力模型上展示任意展开树前缀共享加速的系统。在从SWE-bench任务生成的智能体RL工作负载上,HARTS在多种并行配置下,通过激活重计算实现了4.81至4.87倍的前向/反向/梯度加速,其数值差异与基线自重运行变化相当,且在τ³-Bench训练的前120步中,奖励趋势与基线相似。

英文摘要

Agentic reinforcement learning (RL) often produces irregular rollout trees with shared histories. Training root-to-leaf trajectories independently recomputes these shared prefixes. Existing systems primarily target full-attention models and lack dense, differentiable hybrid-attention execution compatible with activation recomputation. We present HARTS (Hybrid-Attention RL over Tree Structures). HARTS jointly plans microbatches, data-parallel (DP) replica assignments, and microbatch-slot schedules using non-replay compact-token work after prefix compression. For chunkwise linear attention, a linear-time algorithm coordinates chunk-boundary state recovery and replay and produces the minimum number of sequential linear-attention calls under our packed execution model. HARTS preserves the chunkwise state partitioning of trajectory-wise training: it does not repeat projections, MLP/MoE computation, or final outputs, and performs only bounded state replay for numerical alignment. Per round, HARTS batches all branches into one packed call, propagates gradients through differentiable state handoffs, supports activation recomputation, and restores per-token log-probabilities. For deterministic, no-token-drop top-$k$ MoE routing, semantic multiplicities restore MoE-objective token weights and load statistics. Existing RL objectives retain their interface. To our knowledge, HARTS is the first system to demonstrate arbitrary-rollout-tree prefix-sharing speedups on a real hybrid-attention model. On an Agentic RL workload generated from SWE-bench tasks, HARTS achieves $4.81$--$4.87\times$ forward/backward/gradient speedup with activation recomputation across multiple parallel configurations. Its numerical differences are comparable to baseline self-rerun variation, and its reward trend is similar to the baseline over the first 120 steps of $τ^3$-Bench training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑