arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OctoLong:跨仓库代码上下文的训练中优化增强长上下文建模

OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling

Indraneil Paul, Falko Helm, Goran Glavaš, Iryna Gurevych

arXiv 2608.05141首次发表:更新:

AI 中文总结

该研究提出OctoLong流水线,通过跨仓库代码上下文的训练中优化,训练出OctoLong-Instruct长上下文开源LMs,用其替代12%传统语料库可显著提升长上下文相关任务性能及短场景API使用能力。

AI 中文摘要

语言模型(LMs)的上下文长度已大幅提升,这是由上下文学习、自我改进以及长周期智能体工作流的需求推动的。然而,现有的长上下文语料库主要由书籍、学术文章和代码仓库构成,这些资源有限,且在长距离依赖方面往往较为稀缺。在本研究中,我们引入OctoLong,这是一个上下文工程流水线,它利用抽象语法树(AST)解析器、语言服务器后端和包管理器来促进代码引用的递归检索,从而能够整理出包含数百万个标记的、依赖关系丰富的代码上下文。随后,我们训练OctoLong-Instruct,这是一组强大的长上下文开源语言模型(LMs),其基础模型规模从6亿参数到140亿参数不等,训练方式是在约500亿个标记的混合语料库上进行上下文扩展的训练中优化,该混合语料库包含约62亿个OctoLong代码上下文标记,随后进行约100亿个标记的指令调优。我们的训练 ablation(消融实验)以及与18个最先进的开源权重长上下文语言模型的实验评估表明,仅用OctoLong数据替代12%的传统上下文扩展语料库,就能在长距离检索、长期状态跟踪、仓库级代码理解以及下游智能体任务中取得显著提升,同时还能增强短上下文编码场景中的API使用能力。

英文摘要

Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows. Existing long-context corpora, however, are dominated by books, academic articles, and code repositories, which are finite resources and often scarce in long-distance dependencies. In this work, we introduce OctoLong, a context engineering pipeline that instruments an AST parser, a language server backend, and a package manager to facilitate the recursive retrieval of code references, enabling the curation of dependency-rich code contexts of millions of tokens in length. We then train OctoLong-Instruct, a suite of capable long-context open LMs, derived from base models ranging in size from 600M to 14B parameters, via context-extension mid-training on a ~50B-token mixture containing ~6.2B tokens of OctoLong code contexts, followed by ~10B tokens of instruction tuning. Our training ablations and experimental evaluations against 18 state-of-the-art open-weight long-context LMs show that supplanting just 12% of traditional context-extension corpora with OctoLong data yields substantial gains in long-range retrieval, long-term state tracking, repository-level code understanding, and downstream agentic tasks, while also enhancing API usage in short-context coding scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑