发表机构
Peking University; Baidu Inc.; Tsinghua University(北京大学; 百度公司; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出先高亮后总结(H2S)范式,通过先识别证据再生成紧凑摘要来提升长上下文推理,构建数据集与强化学习方法,在基准上以较小模型超越更强基线。
AI 中文摘要
长上下文理解要求大型语言模型(LLMs)对冗长的文档、对话和代码进行推理,然而与任务相关的证据往往稀少且分散在大量无关和冗余的内容之中。我们提出了“先高亮后总结”(H2S)这一压缩后推理范式,它首先识别出基于源文本的、与问题相关的证据,然后将其整合为紧凑的、基于问题的摘要,最后才生成最终答案。为了训练这一行为,我们构建了H2S-Dataset,包含来自11个基准家族的6,647个示例,平均上下文长度为43.9K个词元,并引入了H2S-RL,它在最终答案正确性之外,还为证据选择和摘要构建提供过程级奖励。我们在H2S-Bench(一个包含七项任务的长上下文套件)上进行了评估。在共享的128K输入和4K输出预算下,H2S-14B取得了32.60的平均分,比Qwen3.8-27B高出10.17分,并在所评估的开源模型中获得了最强的总体结果。H2S-14B还取得了最高的证据摘要质量得分,并且在仅使用4K输出预算的情况下,保留了其16K预算性能的97.1%。这些结果表明,显式地选择和整合证据能够改善长上下文推理,同时实现更紧凑的生成。
英文摘要
Long-context understanding requires large language models (LLMs) to reason over lengthy documents, conversations, and code, yet task-relevant evidence is often sparse and scattered amid substantial irrelevant and redundant content. We propose Highlight-Then-Summarize (H2S), a compress-then-reason paradigm that first identifies source-grounded, question-relevant evidence and then integrates it into a compact, question-conditioned summary before producing the final answer. To train this behavior, we construct H2S-Dataset, comprising 6,647 examples from 11 benchmark families with an average context length of 43.9K tokens, and introduce H2S-RL, which provides process-level rewards for evidence selection and summary construction in addition to final-answer correctness. We evaluate on H2S-Bench, a seven-task long-context suite. Under a shared 128K input and 4K output budget, H2S-14B achieves an average score of 32.60, outperforming Qwen3.8-27B by 10.17 points and obtaining the strongest overall result among the evaluated open-source models. H2S-14B also achieves the highest Evidence-Summary Quality score and retains 97.1% of its 16K-budget performance with only a 4K output budget. These results show that explicitly selecting and integrating evidence improves long-context reasoning while enabling more compact generation.
Comments23 pages, 13 figures. Zhaoyuan Xia and Qinghongbing Xie contributed equally. Corresponding authors: Dai Dai, Tong Mo, and Long Zeng. Code and data are available at https://github.com/X-Luffy/Highlight-Then-Summarize