arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

实现真正重要之事:面向大规模推理的原则性上下文表示

Realize What Matters: Principled Context Representation for Large-Scale Reasoning

Michael Theologitis, Dean Light, Shuyue Stella Li, Benjamin Newman, Yulia Tsvetkov, Dan Suciu

arXiv 2609.27173首次发表:更新:

发表机构

University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于相关性实现认知理论的原则性上下文表示方法R3Con,在大型文档推理基准上以较小模型超越大模型,显著提升性能并降低成本。

AI 中文摘要

在科学、医学、法律和金融等领域解决复杂任务,通常需要整合散布在远超模型上下文限制的、庞大且异构的信息源中的相互依赖信息。现有方法通过将信息组织成更易于模型推理的表示形式(如图、文本记忆和检索集合)来应对这一挑战。这些表示决定了下游推理的可能性及其最终成败;然而,其设计和构建在很大程度上仍是临时性的。在本工作中,借鉴相关性实现的认知理论,我们提出了设计能够构建大规模上下文有效表示的AI系统的具体原则。我们分析了现有方法,并展示了其成功与失败如何映射到这些原则的遵循程度上,同时引入了R3Con,一个旨在更系统地落实这些原则的框架。我们在两个近期的大规模文档语料库推理基准上,将R3Con与九个最先进的基线方法进行了评估。在这些基准上,R3Con显著优于最强基线,分别高出20和8.4个百分点。它还使较小的模型能够超越更大的模型:使用4B和9B模型的R3Con优于所有评估的35B基线,而使用35B-A3B模型的R3Con在成本降低3.7倍的情况下超越了使用Claude-Sonnet-5的Claude Code。我们的结果表明,遵循我们原则性方法的上下文表示可以减少对模型规模的依赖,预示着由较小模型驱动的、具有前沿性能的AI系统的未来。我们的代码可在https://github.com/michaeltheologitis/r3con获取。

英文摘要

Solving complex tasks in domains such as science, medicine, law, and finance often requires assembling interdependent information scattered across vast, heterogeneous sources far beyond model context limits. Existing approaches tackle this challenge by organizing information into more manageable representations over which models can reason, such as graphs, textual memories, and retrieval collections. These representations dictate what downstream reasoning is possible and, ultimately, whether it succeeds; yet their design and construction remain largely ad hoc. In this work, drawing on the cognitive theory of relevance realization, we propose concrete principles for designing AI systems that construct effective representations of very large contexts. We analyze existing approaches and show how their successes and failures map onto their alignment with these principles, and introduce R3Con, a harness designed to operationalize the principles more systematically. We evaluate R3Con against nine state-of-the-art baselines on two recent benchmarks of reasoning over large document corpora. On these benchmarks, R3Con substantially outperforms the strongest baseline, by $20$ and $8.4$ percentage points. It also enables smaller models to outperform much larger ones: R3Con with 4B and 9B models outperforms all evaluated 35B baselines, while R3Con with a 35B-A3B model outperforms Claude Code with Claude-Sonnet-5 at $3.7\times$ lower cost. Our results show that context representations following our principled approach can reduce reliance on model scale, pointing toward a future of AI systems with frontier-level performance powered by smaller models. Our code is available at https://github.com/michaeltheologitis/r3con

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑