通过证据链评估进行校准的选择性事实核查
Calibrated Selective Fact-Checking via Evidence Chain Evaluation
浏览论文内容
中文总结 AI 辅助
针对大语言模型事实核查的可靠性问题,提出证据链评估框架ECE,通过网络搜索等收集证据并返回结构化裁决。在ECE - Bench上有较好表现,虽总体校准指标未超最强基线,但实现了选择性预测权衡,弃权可处理弱证据。
中文摘要 AI 辅助
大语言模型在事实核查方面能达到较高准确率,但强制二元决策存在可靠性问题。本文通过证据链评估(ECE)解决此问题,它是一种选择性事实核查框架,允许通过不确定裁决弃权。评估系统是使用工具的验证代理,通过网络搜索等收集证据并返回结构化裁决。在ECE - Bench上,ECE在已回答声明上实现了91.6%的标准准确率、93.7%的覆盖率和97.8%的选择性准确率。虽在总体校准指标上未超最强检索基线,但实现了明确的选择性预测权衡,弃权起到了处理弱证据的安全导向机制作用。
英文摘要
Large language models (LLMs) can achieve strong fact-checking accuracy, yet forced binary decisions conceal a critical reliability problem: systems may issue confident verdicts even when supporting evidence is weak, sparse, or internally inconsistent. We address this issue through Evidence Chain Evaluation (ECE), a selective fact-checking framework that permits abstention via an uncertain verdict instead of requiring a true/false decision for every claim. The evaluated system is a tool-using verification agent that gathers evidence through web search, scholarly search, and executable checks, and then returns a structured verdict with confidence and source-level metadata. On ECE-Bench, ECE achieves 91.6% standard accuracy, 93.7% coverage, and 97.8% selective accuracy on answered claims. Although ECE does not outperform the strongest retrieval baseline on aggregate calibration metrics such as Expected Calibration Error, Brier score, or AURC, it delivers a clear selective-prediction trade-off: the system maintains very high accuracy on answered claims while deferring 6 of 95 cases. These deferred cases are concentrated in lower-reliability evidence settings (5/6 at source level L4), supporting the view that abstention functions as a safety-oriented mechanism for handling epistemically weak evidence. Code is available at https://github.com/ cheshireyang/ECE.git
发表机构
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。