发表机构
Shanghai Jiao Tong University; The Hong Kong University of Science and Technology (Guangzhou)(上海交通大学; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对检索增强模型在证据缺失时仍作答的问题,提出机制边界对齐(RBA)方法,通过匹配变体训练单一阅读器,在支持时作答、缺失时弃权,显著降低不支持作答率并保持支持准确性。
AI 中文摘要
检索增强语言模型本应基于检索到的证据作答,但在实践中,当证据缺失时,它们往往仍会继续作答。我们将这一行为追溯到训练信号上:以答案为中心的微调未对不支持的上下文设定目标,因此无法区分弃权(不执行)的阅读器与猜测的阅读器,并且即使支持的准确性有所提高,不支持的作答率仍接近100%。我们引入了机制边界对齐(Regime Boundary Alignment, RBA),该方法在相同问题和黄金答案的匹配变体上训练单一阅读器。该阅读器被训练为:当上下文支持时(包括同时存在冲突证据的情况)生成黄金答案,而当正确的支持被移除时则弃权(不执行);推理过程为普通解码,无需验证器、阈值或机制标签。在三个多跳问答数据集和三种随机种子上,与冲突聚焦训练相比,RBA将不支持作答率降低了超过六十个百分点,同时保持了相同的支持准确性。在一个留出的TriviaQA检索缺失切片上,同一阅读器将不支持作答率从100%降至1%以下,同时还提高了支持准确性。这些结果表明,证据门控作答必须在支持边界的两侧进行学习。
英文摘要
Retrieval-augmented language models are expected to answer from the retrieved evidence, but in practice they often keep answering when that evidence is missing. We trace this behavior to the training signal: answer-focused fine-tuning assigns no target to unsupported contexts, so it cannot distinguish a reader that abstains from one that guesses, and unsupported answering stays near 100% even as supported accuracy improves. We introduce Regime Boundary Alignment (RBA), which trains a single reader on matched variants of the same question and gold answer. The reader is trained to produce the gold answer when the context supports it, including when conflicting evidence is also present, and to abstain when the correct support is removed; inference is ordinary decoding, with no verifier, threshold, or regime label. On three multi-hop QA datasets across three seeds, RBA reduces the unsupported-answer rate by more than sixty percentage points relative to conflict-focused training while matching its supported accuracy. On a held-out TriviaQA retrieval-miss slice, the same reader reduces unsupported answering from 100% to below 1% while also improving supported accuracy. These results indicate that evidence-gated answering must be learned on both sides of the support boundary.