arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.13683cs.CLcs.AI

通过原生检索增强推理提升上下文保真度

Improving Context Fidelity via Native Retrieval-Augmented Reasoning

  • DIRO, Université de Montréal(蒙特利尔大学计算机科学与运筹学系)
  • MetaGPT
  • Mila - Quebec AI Institute(米拉-魁北克人工智能研究所)
  • McGill University(麦吉尔大学)
  • Yale University(耶鲁大学)
  • Canada CIFAR AI Chair(加拿大CIFAR人工智能教席)

机构由 AI 辅助整理,请以论文原文为准。

Suyuchen Wang, Jinlin Wang, Xinyu Wang, Shiqi Li, Xiangru Tang, Sirui Hong, Xiao-Wen Chang, Chenglin Wu, Bang Liu

更新

AI总结:

提出原生检索增强推理框架CARE,通过模型自身检索能力在推理链中显式整合上下文证据,显著提升了LLMs的上下文保真度及知识密集型任务的准确性。

AI中文摘要:

大型语言模型(LLMs)在上下文保真度方面常面临困难,基于提供的信息回答问题时会产生不一致的答案。现有方法要么依赖昂贵的监督微调在生成答案后产生证据,要么训练模型执行网络搜索,却未必能改善对给定上下文的利用。我们提出CARE,一种新颖的原生检索增强推理框架,教导LLMs利用模型自身的检索能力,在推理过程中显式整合上下文证据。该方法仅需有限的标注证据数据,便可通过推理链中策略性检索的上下文token,显著提升检索准确率和答案生成性能。在多个真实世界和反事实QA基准上的大量实验表明,该方法大幅优于监督微调、传统检索增强生成方法及外部检索解决方案。这项工作在使LLMs对知识密集型任务更准确、可靠和高效方面取得了根本性进展。

英文摘要:

Large language models (LLMs) often struggle with context fidelity, producing inconsistent answers when responding to questions based on provided information. Existing approaches either rely on expensive supervised fine-tuning to generate evidence post-answer or train models to perform web searches without necessarily improving utilization of the given context. We propose CARE, a novel native retrieval-augmented reasoning framework that teaches LLMs to explicitly integrate in-context evidence within their reasoning process with the model's own retrieval capabilities. Our method requires limited labeled evidence data while significantly enhancing both retrieval accuracy and answer generation performance through strategically retrieved in-context tokens in the reasoning chain. Extensive experiments on multiple real-world and counterfactual QA benchmarks demonstrate that our approach substantially outperforms supervised fine-tuning, traditional retrieval-augmented generation methods, and external retrieval solutions. This work represents a fundamental advancement in making LLMs more accurate, reliable, and efficient for knowledge-intensive tasks.

补充信息

↑