arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27661cs.CLcs.AI

先知后答:为可靠检索增强生成(RAG)解码语言模型

Knowing Before Answering: Decoding Language Models for Reliable RAG

  • City St George’s, University of London(伦敦城市圣乔治大学)
  • The Alan Turing Institute(阿兰·图灵研究所)

机构由 AI 辅助整理,请以论文原文为准。

Syed Mahbubul Huq, Christopher Child, Tillman Weyde, Pranava Madhyastha

AI总结:

本研究针对RAG中检索信息不足或冲突的问题,构建标注数据集,利用语言模型的隐藏激活等特征训练线性分类器,实现可靠的RAG分类,性能优于基线及专用RAG模型。

AI中文摘要:

在检索增强生成(RAG)中,检索到的信息可能不足以回答问题,也可能包含相互冲突的信息。系统不仅要知道何时可以回答问题,还需能识别RAG提供的文档信息不足或存在冲突的情况。该问题可被建模为一个三分类问题,即利用模型的内部信号判断输入中的检索信息属于充足、不足还是冲突三类。我们构建了一个受控基准数据集,该数据集模拟RAG场景并包含虚构信息,每个实例被标注为可回答、信息不足或信息冲突。我们将隐藏激活值和注意力衍生特征作为输入,训练一个轻量级线性模型来区分这三类情况。在涵盖不同架构和规模的16个语言模型上,我们的基于特征的路由器始终优于基于提示的基线方法以及专用RAG模型的性能。我们还对模型的信息动态进行了分析,结果表明,分类所需的最具信息量的信号出现在模型的中间层,在大多数测试模型中,隐藏激活状态比注意力值或MLP特征输出更有效。总体而言,我们的结果表明语言模型内部编码了检索到的证据是否足以支持回答问题,且该信号可被可靠解码用于RAG的分类处理。

英文摘要:

In Retrieval-Augmented Generation (RAG), retrieval may provide insufficient or conflicting information needed to answer a question. The system should not only know when to answer but also be able to identify cases in which the documents provided in RAG are insufficient or contain conflicting information. This can be framed as a three-way classification problem, where we use the model's internal signals to determine whether the provided information in the input can be classified as sufficient, insufficient, or conflicting. We create a controlled benchmark dataset that replicates a RAG setup with fictitious information and labels each instance as answerable, insufficient, or conflicting. We use hidden activations and attention-derived features as inputs to train a lightweight linear model to distinguish among the three classes. Across 16 language models spanning different architectures and a range of model sizes, our feature-based router consistently outperforms prompting-based baselines and the performance of specialised RAG-models. We further conduct analyses into the information dynamics of the models. We show that the most informative signals for the classification are available in the middle layers, with hidden activation states being more effective than attention values or the MLP-feature outputs in most of the tested models. Overall, our results suggest that language models internally encode whether retrieved evidence is sufficient to support answering, and that this signal can be decoded reliably for RAG triage.

补充信息

↑