arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.06160cs.CLcs.AI

LongCrafter:通过证据图引导的指令合成实现多样化的长上下文理解

LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis

  • University of Chinese Academy of Sciences(中国科学院大学)
  • The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所认知与决策智能复杂系统重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

Chenhao Yuan, Yinhao Xu, Shuwen Xu, Xizhi Yang, Jiaxiang Liu, Chenxi Zhou, Shaoping Huang, Haolin Ren, Pengfei Cao, Jun Zhao, Kang Liu

AI总结:

针对现有长上下文理解方法的局限,提出LongCrafter框架,结合分层任务分类法与证据管道,生成多样化长上下文SFT数据,微调后的模型在多个数据集上表现优异,能有效缓解“中间迷失”问题。

AI中文摘要:

合成长上下文监督微调(SFT)数据是增强大语言模型(LLMs)长上下文理解的可扩展方法,但现有方法存在任务覆盖窄、指令难度不足和缺乏忠实监督三个局限。我们提出LongCrafter,一个将分层任务分类法与基于证据的管道相结合的结构化合成框架。该分类法将长上下文理解组织成本地/浅层和全局/深层,并产生32种细粒度任务类型作为全局生成先验。在LongCrafter数据上微调的模型在多个数据集上优于所有SFT基线甚至官方后训练模型,在高难度任务上增益最大。进一步分析表明,LongCrafter数据更多样化且在难度级别上分布更好,训练后的模型能稳健定位证据,有效缓解“中间迷失”问题。

英文摘要:

Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instruction difficulty, and a lack of faithfulness supervision. We propose \textbf{LongCrafter}, a structured synthesis framework that couples a hierarchical task taxonomy with an evidence-grounded pipeline. The taxonomy organizes long-context understanding into local/shallow and global/deep levels and yields 32 fine-grained task types that serve as a global generative prior. Guided by this taxonomy, LongCrafter constructs task-aligned long contexts, decomposes them into explicit evidence graphs that model cross-paragraph dependencies, and generates instruction--response pairs strictly grounded in the located evidence spans, ensuring both controllable difficulty and faithful, traceable reasoning. Models fine-tuned on LongCrafter data outperform all SFT baselines and even the official post-trained models on LongBench, LongBench~v2, and LooGLE across both Qwen2.5-7B and LLaMA-3.1-8B, with the largest gains on high-difficulty tasks. Further analysis shows that LongCrafter data is more diverse and better spread across difficulty levels, and that the trained models locate evidence robustly regardless of position, effectively mitigating the ``lost in the middle'' problem.

↑