发表机构
Université de Toulouse - IRIT UMR 5505; Airbus Protect(图卢兹大学; 空客保护)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对检索增强生成中引用片段不忠实的问题,提出约束混合解码(CHyD),通过硬约束确保输出引用逐字来自检索文档,在技术领域实现近乎完美的提取忠实性。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用作信息检索的接口,但它们仍然容易出现幻觉和忠实性错误,即生成的答案与检索到的证据不一致。虽然检索增强生成(RAG)和最近的混合或半抽取式方法缓解了这一问题,但它们并不能保证引用的或提取的片段是逐字来自检索到的上下文。这种限制在安全关键领域可能造成严重后果,在这些领域中,答案必须与认证文档完全匹配。我们引入了约束混合解码(CHyD),这是一种用于推测性RAG的新颖的忠实性优先范式。虽然传统的推测性解码针对推理速度进行了优化,但CHyD重新利用了这种架构,以确保在提取模式被正确触发时进行忠实的逐字证据提取。我们的方法强制执行硬解码约束,将生成限制在检索文档中存在的连续片段上。这种设计提供了一个稳健但直接的保证:输出中任何明确引用的片段都逐字出现在提供的上下文中。我们在最先进的LLMs上,对多种抽象式、抽取式和半抽取式问答基准进行了评估,包括受飞机维护启发的技术数据集。结果表明,现有的混合方法经常幻觉引用的片段,在技术领域中精确提取准确率降至40%以下。相比之下,我们的方法无论使用何种模型,都实现了近乎完美的提取忠实性。虽然强制执行硬约束引入了与流畅性相关指标的权衡,但我们的方法提高了精确答案的正确性,并总体上保持竞争力,突显了其在安全关键信息检索应用中的适用性。
英文摘要
Large Language Models (LLMs) are increasingly used as interfaces for information retrieval, but they remain prone to hallucinations and faithfulness errors, in which the generated answers diverge from the retrieved evidence. While Retrieval-Augmented Generation (RAG) and recent hybrid or semi-extractive approaches mitigate this issue, they do not guarantee that quoted or extracted spans are verbatim from the retrieved context. This limitation can have severe consequences in safety-critical domains, where answers must exactly match certified documentation. We introduce Constrained Hybrid Decoding (CHyD), a novel faithfulness-first paradigm for speculative RAG. While traditional speculative decoding is optimized for inference speed, CHyD repurposes this architecture to ensure faithful verbatim evidence extraction when the extraction mode is correctly triggered. Our approach enforces hard decoding constraints that restrict generation to continuous spans present in the retrieved documents. This design provides a robust but straightforward guarantee: any explicitly quoted span in the output appears verbatim in the provided context. We evaluate our method across state-of-the-art LLMs on diverse abstractive, extractive, and semi-extractive QA benchmarks, including technical datasets motivated by aircraft maintenance. Results show that existing hybrid methods frequently hallucinate quoted spans, with exact extraction accuracy dropping below 40% in technical domains. In contrast, our approach achieves near-perfect extraction faithfulness regardless of the model used. Although enforcing hard constraints introduces a trade-off with fluency-oriented metrics, our method improves exact answer correctness and remains competitive overall, highlighting its suitability for safety-critical information retrieval applications.