发表机构
The Hong Kong University of Science and Technology (Guangzhou); Hunyuan Team, Tencent; Peking University(香港科技大学(广州); 腾讯混元团队; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出TrajLong框架,将智能体轨迹编译为长上下文训练任务,通过密集监督增强三种原子能力,在多个基准上验证了性能提升。
AI 中文摘要
用于编码、搜索和工作场所任务的LLM智能体日益依赖长上下文能力,以有效聚合和推理扩展的交互历史。近期工作已将智能体轨迹纳入中期训练阶段,利用其天然的长且交互丰富的结构。然而,如何将这些轨迹中的信息组织成有效的中期训练监督仍未被充分探索。在本工作中,我们研究了长上下文与智能体原子能力之间的关系,并引入了TrajLong,一种新颖的框架,将轨迹编译为具有密集监督的长上下文训练任务,针对三种代表性的原子能力:证据基础、跨证据聚合和时间状态维护。我们使用TrajLong编译的数据对Qwen3-14B-Base和Qwen3-30B-A3B-Base进行中期训练,随后进行监督微调。在6个长上下文和12个智能体基准上的实验展示了广泛的性能提升,受控消融实验显示相对于原始和掩码轨迹基线的改进。能力级分析进一步揭示了长上下文与智能体原子能力之间的任务依赖性关联。这些发现表明,长上下文推理和智能体执行的共享能力需求为设计中期训练数据以发展下游智能体能力提供了原则性基础。
英文摘要
LLM agents for coding, search, and workplace tasks increasingly rely on long-context capabilities to effectively aggregate and reason over extended interaction histories. Recent work has incorporated agent trajectories into mid-training stage, drawing on their naturally long and interaction-rich structure. Yet how to organize the information within these trajectories into effective mid-training supervision remains underexplored. In this work, we investigate the relationship between long-context and agent atomic capabilities and introduce TrajLong, a novel framework that compiles trajectories into long-context training tasks with dense supervision, targeting three representative atomic capabilities: evidence grounding, cross-evidence aggregation, and temporal state maintenance. We mid-train Qwen3-14B-Base and Qwen3-30B-A3B-Base with data compiled by TrajLong, followed by supervised fine-tuning. Experiments on 6 long-context and 12 agent benchmarks demonstrate broad performance gains, with controlled ablations showing improvements over raw and masked trajectory baselines. Capability-level analyses further reveal task-dependent associations between long-context and agent atomic capabilities. These findings suggest that the shared capability demands of long-context reasoning and agent execution provide a principled basis for designing mid-training data to develop downstream agent capabilities.