选择长视界轨迹以实现可靠高效的终端智能体训练
Selecting Long-Horizon Trajectories for Reliable and Efficient Terminal-Agent Training
浏览论文内容
中文总结 AI 辅助
针对终端智能体训练中监督视界的选择问题,提出选择性长视界细化方法,通过先短后长并筛选高概率延续,提升任务成功率并降低训练成本。
中文摘要 AI 辅助
终端智能体通常通过模仿长的教师轨迹进行训练,然而每条轨迹中应监督多少内容仍未得到探索。我们研究了“监督视界”,即训练中保留的轨迹令牌数量,并表明它是影响可靠性和成本的关键设计轴。可靠性随着更长的视界而提高,但会饱和:在Terminal-Bench上,12K令牌的视界比16K解决了更多任务($29\pm0.7$ vs. $26\pm0.8$),同时训练时间减少30%。视界也塑造了智能体行为:短视界导致过早终止,中等视界产生有效的错误恢复,长视界则引发过度坚持。我们通过偏差-复杂度界分析了这种饱和现象,其中更长的监督减少了时间监督偏差,但增加了来自更异质的后期历史的有限样本估计误差。在此分析的指导下,我们提出了“选择性长视界细化”,该方法首先在短前缀上训练,然后仅在热启动模型下最有可能的延续上进行细化。它始终优于完整的长视界训练。在16K时,它将成功尝试从$110\pm2.7$提高到$126\pm2.1$,将八次尝试中至少六次成功的任务从$9\pm0.7$提高到$14\pm0.6$;使用一半的长视界数据,它仍达到$122\pm2.4$,同时训练时间减少23%。这些收益跨基准迁移,在Terminal-Bench v2.0上从$64\pm2.6$提高到$73\pm2.1$,在OpenThoughts-TBLite上从$137\pm2.7$提高到$155\pm2.2$。对于长视界监督,选择正确的轨迹比训练所有轨迹更重要。
英文摘要
Terminal agents are commonly trained by imitating long teacher trajectories, yet how much of each trajectory to supervise remains unexplored. We study the \emph{supervision horizon}, the number of trajectory tokens retained for training, and show that it is a key design axis for reliability and cost. Reliability improves with longer horizons but saturates: on Terminal-Bench, a 12K-token horizon solves more tasks than 16K ($29\pm0.7$ vs.\ $26\pm0.8$) while requiring 30\% less training time. The horizon also shapes agent behavior: short horizons cause premature termination, intermediate horizons yield productive error recovery, and long horizons induce over-persistence. We analyze this saturation through a bias--complexity bound, in which longer supervision reduces temporal supervision bias but increases finite-sample estimation error from more heterogeneous late-stage histories. Guided by this analysis, we propose \emph{selective long-horizon refinement}, which first trains on short prefixes and then refines only on continuations that are most likely under the warm-start model. It consistently outperforms full long-horizon training. At 16K, it raises successful attempts from $110\pm2.7$ to $126\pm2.1$ and tasks solved in at least six of eight attempts from $9\pm0.7$ to $14\pm0.6$; with half of the long-horizon data, it still reaches $122\pm2.4$ while cutting training time by 23\%. The gains transfer across benchmarks, from $64\pm2.6$ to $73\pm2.1$ on Terminal-Bench v2.0 and from $137\pm2.7$ to $155\pm2.2$ on OpenThoughts-TBLite. For long-horizon supervision, selecting the right trajectories matters more than training on all of them.
发表机构
- Virginia Tech(弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。