arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08943cs.CLcs.AIcs.IR

评估与改进基于多轮证据消融的LLM证据接地事实核查

Evaluating and Improving Evidence-Grounded Fact-Checking in LLMs via Multi-Round Evidence Ablation

Xingyu Deng, Mingzi Cao, Nikolaos Aletras, Xi Wang, Mark Stevenson

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出FAE评估框架与REAL训练方法,发现LLM事实核查过度依赖参数知识,通过反事实证据监督训练增强其证据依赖能力,在四个数据集上优于标准微调模型。

中文摘要 AI 辅助

自动事实核查系统根据相关文档中的证据来评估声明的真实性。大型语言模型(LLMs)凭借其通用推理能力,在事实核查中展现出强劲性能。然而,它们是否忠实地利用所提供的证据来达成真实性判断,还是依赖于参数化知识,仍不清楚。为探究此问题,我们引入了事实消融评估(FAE),一种新的评估框架,通过迭代消融引用的证据来评估LLMs是否相应地修正其预测。我们的实证结果表明,当前现成的LLMs作为事实核查系统,更多地依赖其参数化知识而非所提供的证据。为弥合预测准确性与证据接地之间的差距,我们提出了REAL(严格证据消融学习),一种训练框架,通过反事实证据监督促进LLM作为验证器模型的证据依赖型验证。在四个不同领域的事实核查数据集上的实验表明,与标准微调模型相比,使用REAL训练的模型获得了更优越的证据依赖能力。我们的发现强调,强事实核查性能仍可与弱证据依赖共存,而REAL鼓励真实性预测与支持性证据的可用性保持更紧密的联系。

英文摘要

Automatic fact-checking systems assess the veracity of claims given evidence from relevant documents. Large Language Models (LLMs) have demonstrated strong performance in fact-checking due to their general reasoning capabilities. However, it remains unclear whether they faithfully make use of the evidence provided to reach veracity judgments or rely on parametric knowledge. To investigate this, we introduce Fact-Ablated Evaluation (FAE), a new evaluation framework that iteratively ablates the cited evidence to assess whether LLMs revise their predictions accordingly. Our empirical results show that current off-the-shelf LLMs as fact-checking systems rely more on their parametric knowledge than on the evidence provided. To bridge this gap between prediction accuracy and evidence grounding, we propose REAL (Rigorous Evidence Ablation Learning), a training framework that promotes evidence-dependent verification through counterfactual evidence supervision for the LLM-as-verifier models. Experiments on four fact-checking datasets across different domains demonstrate that models trained with REAL obtain superior evidence-dependent capabilities compared to standard fine-tuned models. Our findings highlight that strong fact-checking performance can still coexist with weak evidence dependency, while REAL encourages veracity predictions to remain more closely tied to the availability of supporting evidence.

发表机构

  • University of Sheffield(谢菲尔德大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑