arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18240cs.AI

通过证据链评估进行校准的选择性事实核查

Calibrated Selective Fact-Checking via Evidence Chain Evaluation

Dekun Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型事实核查的可靠性问题,提出证据链评估框架ECE,通过网络搜索等收集证据并返回结构化裁决。在ECE - Bench上有较好表现,虽总体校准指标未超最强基线,但实现了选择性预测权衡,弃权可处理弱证据。

中文摘要 AI 辅助

大语言模型在事实核查方面能达到较高准确率,但强制二元决策存在可靠性问题。本文通过证据链评估(ECE)解决此问题,它是一种选择性事实核查框架,允许通过不确定裁决弃权。评估系统是使用工具的验证代理,通过网络搜索等收集证据并返回结构化裁决。在ECE - Bench上,ECE在已回答声明上实现了91.6%的标准准确率、93.7%的覆盖率和97.8%的选择性准确率。虽在总体校准指标上未超最强检索基线,但实现了明确的选择性预测权衡,弃权起到了处理弱证据的安全导向机制作用。

英文摘要

Large language models (LLMs) can achieve strong fact-checking accuracy, yet forced binary decisions conceal a critical reliability problem: systems may issue confident verdicts even when supporting evidence is weak, sparse, or internally inconsistent. We address this issue through Evidence Chain Evaluation (ECE), a selective fact-checking framework that permits abstention via an uncertain verdict instead of requiring a true/false decision for every claim. The evaluated system is a tool-using verification agent that gathers evidence through web search, scholarly search, and executable checks, and then returns a structured verdict with confidence and source-level metadata. On ECE-Bench, ECE achieves 91.6% standard accuracy, 93.7% coverage, and 97.8% selective accuracy on answered claims. Although ECE does not outperform the strongest retrieval baseline on aggregate calibration metrics such as Expected Calibration Error, Brier score, or AURC, it delivers a clear selective-prediction trade-off: the system maintains very high accuracy on answered claims while deferring 6 of 95 cases. These deferred cases are concentrated in lower-reliability evidence settings (5/6 at source level L4), supporting the view that abstention functions as a safety-oriented mechanism for handling epistemically weak evidence. Code is available at https://github.com/ cheshireyang/ECE.git

发表机构

  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

↑