发表机构
University of Waterloo; University of British Columbia; NVIDIA; Verdent AI; Vector Institute(滑铁卢大学; 英属哥伦比亚大学; 英伟达公司; Verdent人工智能公司; 向量研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对编码智能体将外部工具返回集成到推理中的问题,利用函数感知中间填充进行中间训练,在多个模型上提升了性能,并减轻了后训练对非智能体编码等基准测试的能力侵蚀。
AI 中文摘要
编码智能体必须将外部工具返回结果集成到正在进行的推理中,而标准的代码从左到右预训练仅在正向暴露此能力。我们观察到编码智能体的动作-观察-延续循环在结构上与函数调用站点同构。我们通过函数感知中间填充(FIM)中间训练来利用这一点,这是一种自监督目标,通过程序依赖图分析和复杂性可推断性双重标准来屏蔽函数。我们在从968个GitHub仓库抽取的26亿令牌的净化语料库上对Qwen2.5-Coder-Instruct(7B/14B)和Qwen3-8B进行中间训练,然后应用现有的智能体后训练管道。中间训练在不同模型和后训练管道上都有提升,还减轻了后训练对非智能体编码和非编码工具使用基准测试的能力侵蚀。
英文摘要
Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the action-observation-continuation loop of a coding agent is structurally isomorphic to a function call site, where a caller binds arguments, a callee returns a value computed elsewhere, and downstream code consumes that value. This conditioning structure exists at internet scale in ordinary code. We exploit it through function-aware fill-in-the-middle (FIM) mid-training: a self-supervised objective that masks functions selected via program dependency graph analysis and a complexity-inferability double criterion. We mid-train Qwen2.5-Coder-Instruct (7B/14B) and Qwen3-8B on a 2.6B-token decontaminated corpus drawn from 968 GitHub repositories, then apply existing agentic post-training pipelines. Mid-training improves SWE-Bench-Verified by +2.8/+3.0 at 7B/14B and by +3.2 on Qwen3-8B; SWE-Bench-Lite gains are +3.7/+4.0/+5.4 on the same models. The improvement holds across two post-training pipelines (R2E-Gym, SWE-Smith) and on a non-Qwen2.5 base (Qwen3-8B with SWE-Lego). Beyond in-domain gains, mid-training also mitigates the capability erosion that agentic post-training otherwise inflicts on non-agent coding (e.g., LiveCodeBench) and non-coding tool-use benchmarks (tau-bench, BFCL): although the mid-training corpus contains Python code only, the function-call inductive bias survives post-training and yields consistent gains.