arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23514cs.CLcs.AIcs.MM

新主张还是似曾相识?重新思考多模态自动事实核查的“无污染”动态评估

Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking

Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, Yupeng Li

首次发表
浏览论文内容

中文总结 AI 辅助

研究多模态自动事实核查中基准的污染风险,通过实证研究静态和动态基准,发现动态评估虽能减少但不能消除污染,新主张可多种方式验证,污染会致性能膨胀和排名扭曲,为可信评估提供实用指南。

中文摘要 AI 辅助

多模态自动事实核查(MAFC)通过检索和推理外部证据来验证主张。然而,大多数现有静态基准存在污染风险:主要由可利用大语言模型内部知识而非外部证据验证的过时主张组成,这会夸大性能估计且无法反映对新主张的真实能力。为解决此问题,新兴动态基准收集大语言模型知识截止日期后发布的主张并假定无污染。本文通过实证研究最先进的静态AVeriTeC基准和新构建的动态ClaimReview2025Q4基准中的污染风险及其对MAFC评估的影响来重新审视这一假设。实验得出16个发现,突出了三个关键结果:动态评估减少但未消除污染风险;许多新发布主张可直接或通过合成截止前的多条公共知识来验证;污染会导致MAFC性能在统计上显著膨胀,扭曲系统排名。鉴于这些发现,在严格污染控制设置下重新评估了最先进的大语言模型。本研究为可信的MAFC评估提供了实用指南。

英文摘要

Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without external evidence. This can inflate performance estimates and fail to reflect true capability on novel claims that require up-to-date information. To address this, emerging dynamic benchmarks collect claims published after LLMs' knowledge cut-off dates, assuming they are uncontaminated. This work revisits this assumption by empirically studying contamination risks in both the state-of-the-art (SOTA) static AVeriTeC benchmark and our newly constructed dynamic ClaimReview2025Q4 benchmark, as well as their impact on MAFC evaluation. Our experiments yield 16 findings, highlighting three key results: (1) Dynamic evaluation reduces but does not eliminate contamination risks, as 17.09\%--29.30\% of post-cut-off claims remain potentially contaminated; (2) Many newly published claims can be verified either directly or by synthesizing multiple pieces of public knowledge available before the cut-off; and (3) Contamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distorting system rankings. In light of these findings, we re-evaluate SOTA LLMs under a strictly contamination-controlled setting. Our study provides practical guidelines for trustworthy MAFC evaluation.

发表机构

  • Hong Kong Baptist University(香港浸会大学)
  • The University of Hong Kong(香港大学)
  • Beijing Normal-Hong Kong Baptist University(北京师范大学-香港浸会大学联合国际学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑