发表机构
Leuphana University of Lüneburg; MBZUAI; University of Zurich(吕讷堡勒芬纳大学; 穆罕默德·本·扎耶德人工智能大学; 苏黎世大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究对九种模型在四个跨领域数据集上开展自动事实核查系统的跨基准评估,发现系统性能依赖领域与指标,检索是主要瓶颈,并发布资源支持可复现研究。
AI 中文摘要
自动事实核查(AFC)系统会检索证据并预测主张的真实性,但现有评估忽略了简单基线,且这些系统仅针对单一基准开发,无法跨领域泛化。此前尚无研究在不同数据集上对完整的“先检索后验证”两阶段流程进行交叉评估,仅存在仅检索研究(Thakur等人,2021)和单阶段基准研究(Calamai等人,2025)。我们在涵盖科学、开放网络和气候领域的四个数据集上,对九种模型进行基准测试,包括随机基线、稀疏基线、微调后的Transformer、零样本大语言模型(LLM)以及AVeriTeC 2025共享任务中排名最高的两个系统。研究得出三项关键发现:(1)在ClimateCheck数据集上,仅主张模型和微调模型的表现优于零样本LLM和AVeriTeC 2025顶级系统,表明嘈杂证据会降低真实性预测性能;(2)系统排名高度依赖领域和评估指标:SciFact上表现最佳的模型(宏F1值0.70)在ClimateCheck上降至0.31,而AVeriTeC 2025的冠亚军在不同评估指标和数据集上排名互换;(3)用黄金标注替换检索到的证据可使各模型的真实性准确率提升14-22个百分点,证实检索仍是主要瓶颈。我们发布代码、预处理后的数据集及所有结果,以支持可复现的AFC研究。
英文摘要
Automated fact-checking (AFC) systems retrieve evidence and predict claim veracity, yet evaluations omit simple baselines, systems are developed for a single benchmark and cannot be trusted to generalise across domains. No prior work cross-evaluates the full two-stage retrieve-then-verify pipeline across diverse datasets, complementing retrieval-only studies (Thakur et al., 2021) and single-stage benchmarking studies (Calamai et al., 2025). We benchmark nine models, ranging from random and sparse baselines to fine-tuned transformers, zero-shot LLMs, and the two highest-ranked systems from the AVeriTeC 2025 shared task, across four datasets spanning scientific, open-web, and climate domains. Three findings stand out: (1) on ClimateCheck claim-only and fine-tuned models outperform zero-shot LLM and top-performing AVeriTeC 2025 systems, highlighting that noisy evidence can degrade veracity prediction; (2) system rankings are strongly domain- and metric-dependent: the best model on SciFact (macro-F1 0.70) drops to 0.31 on ClimateCheck, while the AVeriTeC 2025 winner and runner-up swap rankings based on evaluation metrics and datasets; (3) replacing retrieved evidence with gold annotations improves veracity accuracy by 14-22 points across models, confirming retrieval remains primary bottleneck. We release code, pre-processed datasets, and all results to support reproducible AFC research.