arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11955cs.CLcs.LG

R2VC:基于检索、验证与置信度校准的模块化事实验证

R2VC: Modular Fact-Checking with Retrieval, Verification, and Confidence Calibration

  • Stevens Institute of Technology(史蒂文斯理工学院)
  • University of North Carolina(北卡罗来纳大学)

机构由 AI 辅助整理,请以论文原文为准。

Dhruv Dixit, Paritosh Pandey

AI总结:

针对端到端事实验证中证据检索、推理与置信度估计纠缠的问题,提出模块化架构R2VC,结合混合检索、生成候选、NLI验证和置信度校准,在FEVER上提升准确率13.74%,并增强置信度可靠性。

AI中文摘要:

大型语言模型越来越多地被用于自动化事实验证,但端到端的提示方法常常将证据检索、推理和不确定性估计纠缠在一起,使得失败难以诊断,置信度难以信任。我们提出了R2VC,一种模块化的检索、推理、验证、校准架构,用于带有引用和弃权(不执行)功能的基于证据的事实验证。R2VC结合了在维基百科上的混合稀疏+稠密检索、一个经过监督微调和DPO对齐的生成器(可产生多样化的结构化判定候选)、一个用于基于证据的候选选择的外部NLI交叉编码器,以及一个用于置信度估计和选择性弃权(不执行)的轻量级序列级校准器。在FEVER上,采用R2VC的8B骨干模型比基线准确率高出13.74%。消融研究表明,基于验证器的候选选择和置信度校准是性能提升的最大贡献者。移除候选选择使FEVER准确率降至76.24%,而移除校准使Brier分数几乎翻倍至0.161。对250个错误的人工分析进一步表明,检索失败,尤其是错误实体的证据,仍然是主要的瓶颈。总之,这些结果表明,模块化的事实验证流程可以显著提高开放域验证中的预测准确性和置信度可靠性。

英文摘要:

Large language models are increasingly used for automated fact checking, but end-to-end prompting often entangles evidence retrieval, reasoning, and uncertainty estimation, making failures difficult to diagnose and confidence difficult to trust. We present R2VC, a modular retrieve, reason, verify, calibrate architecture for evidence-grounded fact checking with citations and abstention. R2VC combines hybrid sparse+dense retrieval over Wikipedia, a supervised fine-tuned and DPO-aligned generator that produces diverse structured verdict candidates, an external NLI cross-encoder for evidence-based candidate selection, and a lightweight sequence-level calibrator for confidence estimation and selective abstention. On FEVER, an 8B backbone with R2VC achieves 13.74% higher accuracy than baseline. Ablation studies show that verifier-based candidate selection and confidence calibration are the largest contributors to performance. Removing candidate selection drops FEVER accuracy to 76.24%, while removing calibration nearly doubles the Brier score to 0.161. A manual analysis of 250 errors further shows that retrieval failures, especially wrong-entity evidence, remain the dominant bottleneck. Together, these results show that modular fact-checking pipelines can substantially improve both predictive accuracy and confidence reliability in open-domain verification.

补充信息

↑