T1:面向长时程任务的终端智能体强化学习
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
浏览论文内容
中文总结 AI 辅助
T1通过强化学习训练122B混合专家模型,在云沙箱中操作shell执行长时程终端任务,采用TITO和R3稳定优化,在Terminal-Bench 2.1上解决率达64.0%,超越GPT-5.4和GLM-5.1。
中文摘要 AI 辅助
智能体的使用正转向长时程任务,如编程和科学发现,其中终端任务尤为重要。我们提出了T1,一个总参数为122B的混合专家模型,通过强化学习训练,在云沙箱中操作真实shell,每个任务最多可进行300多次工具调用轮次,并通过执行每个任务自身的验证器来获得奖励。我们提供了一套全面的方案:首先,采用激进的热启动以稳定actor-critic训练,通过密集的过程奖励,根据通过验证器的绝对数量对轨迹进行评分。其次,通过TITO构造实现稳定优化,在精确采样的token标识符上训练,并在轮次边界进行漂移修复,以及通过rollout路由重放,记录采样器在每个MoE层的逐token专家选择,并在训练期间重放它们。第三,完全分布外训练语料:与Terminal-Bench 2.1不相交的孤立种子和合成任务,确保性能提升反映的是真实能力迁移而非基准过拟合。综合来看,TITO和R3将训练到推理的对数概率差异从0.021降至0.013,且在损失区域实现了精确对齐的零token漂移。在Terminal-Bench 2.1上,我们的后训练流程将基础模型从43.8%提升至T1的64.0%解决率。在Long-Horizon Terminal Bench上,T1达到27.9%,超越了GPT-5.4和GLM-5.1。
英文摘要
Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier. We provide a comprehensive recipe: First, an aggressively warm-started to stabilize actor-critic training, with a dense process reward scoring trajectories by the absolute number of passing verifiers. Second, stable optimization through TITO construction, training on the exact sampled token identifiers with drift repair at turn boundaries, and rollout routing replay, recording the sampler's per-token expert choices at every MoE layer and replaying them during training. Third, fully out-of-distribution training corpus: isolated seeds and synthesized tasks disjoint from Terminal-Bench 2.1 ensures gains reflect genuine capability transfer over benchmark overfitting. Together, TITO and R3 cut the training-to-inference log-probability difference from 0.021 to 0.013, with exactly aligned zero token drift in the loss region. On Terminal-Bench 2.1, our post-train pipeline raises initial base model from 43.8% to T1 with 64.0% resolved. On Long-Horizon Terminal Bench, T1 reaches 27.9% and surpasses GPT-5.4 and GLM-5.1.
发表机构
- Tencent Hy Foundation Model Frontier(腾讯Hy基础模型前沿)
- National University of Singapore(新加坡国立大学)
- University of Georgia(佐治亚大学)
- Indiana University(印第安纳大学)
- University of Maryland, College Park(马里兰大学帕克分校)
机构由 AI 辅助整理,请以论文原文为准。