RAG引用中的可归因事后合理化:一项受控复现与RLVR比较研究
Attributable Post-Rationalization in RAG Citations: A Controlled Reproduction and an RLVR Comparison
浏览论文内容
中文总结 AI 辅助
本研究通过受控复现和RLVR比较,发现RAG系统普遍存在事后合理化导致的不忠实引用,且RLVR训练无法改善引用忠实性,需单独训练和评估。
中文摘要 AI 辅助
RAG系统可以给出正确的答案,却引用一个它实际并未使用的来源。模型通过事后合理化输出这些不忠实的引用:它们先写出答案,然后将引用附加到任何看起来足够接近的段落上。搜索智能体现在通过可验证奖励的强化学习(RLVR)进行训练,该训练因答案正确而给予奖励。我们询问这种训练是否也教会它们诚实地引用。通过改进现有方法并加入必要的控制,我们比较了一个指令微调模型与从该模型训练出的三个RLVR智能体,在四个问答数据集上的表现,仅使用免费的Kaggle GPU。事后合理化无处不在:在基于维基百科的问题中,大约每七个引用中就有一个是不忠实的。RLVR并未修复这一问题。这些智能体以它们基础模型的比率进行事后合理化,其中一个的表现略差。奖励正确答案在引用忠实性方面毫无收获,因此忠实性必须单独进行训练和衡量。
英文摘要
A RAG system can hand you the right answer and cite a source it did not actually use. Models output these unfaithful citations via post-rationalization: they write the answer first and then attach a citation to whatever passage looks close enough. Search agents are now trained with reinforcement learning from verifiable rewards (RLVR), which pays them for getting the answer right. We asked whether that training also teaches them to cite honestly. Improving an existing methodology with a required control, we compared an instruction-tuned model against three RLVR agents trained from it, on four question-answering datasets, using only free-tier Kaggle GPUs. Post-rationalization is everywhere: on Wikipedia-based questions roughly one citation in seven is unfaithful. RLVR does not fix it. The agents post-rationalize at their base model's rate, and one lands slightly worse. Rewarding correct answers buys nothing in citation faithfulness, so faithfulness has to be trained and measured on its own terms.
发表机构
- Bangladesh University of Engineering and Technology(孟加拉国工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。