发表机构
The University of Sydney(悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对短视频事实核查的证据选择问题,构建MiniVer-V基准,提出两层核查框架与充分性驱动的贪心搜索方法,在减少证据使用量的同时保持核查性能,提升了对证据不足案例的识别能力。
AI 中文摘要
短视频事实核查的核心挑战在于识别哪些证据足以支撑核查结论。现有方法要么向核查器提供所有可用证据,引入噪声;要么通过主题相关性选择证据,将相关性与充分性混为一谈。我们将证据充分性确定为选择标准:即证据子集是否足以支撑可信裁决且无冗余。我们推出MiniVer-V基准,包含195条短视频,带有三类裁决标注(支持、反驳、证据不足),以及5510个多模态证据单元,涵盖视觉关键帧、语音转录文本和网络检索外部来源。我们提出两层核查框架,将从内部证据评估的主张-视频一致性与需额外外部佐证的事实裁决确定相分离。在此基础上,一种充分性驱动的贪心搜索会逐步组合证据,直至达到充分性阈值;当候选池耗尽时,输出“证据不足”,而非强制给出裁决。使用Claude Sonnet 4时,该方法平均使用4.5个证据单元(占全部证据集的16%),达到Macro-F1值0.510,与全证据基线(使用27.7个证据单元,Macro-F1值0.518)在统计上无差异,同时相比不采用弃权(不执行)的相同搜索,显著提升了对证据不足案例的识别能力。该效率结果在GPT-5.5上可复现,而在开放权重的Qwen2.5-72B核查器上仅部分成立。消融实验表明,外部证据对事实裁决不可或缺,而内部视频证据为裁决提供主张-视频一致性的依据。这些发现表明,实现证据高效的核查是可行的,且当证据确实不足时,需要明确的弃权(不执行)机制。
英文摘要
A core challenge in short-video fact-checking is identifying which evidence is sufficient to support a verification conclusion. Existing approaches either give the verifier all available evidence, introducing noise, or select evidence by topical relevance, which conflates relatedness with sufficiency. We identify evidential sufficiency as the selection criterion: whether a subset of evidence is adequate to support a confident verdict without redundancy. We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources. We propose a two-layer verification framework that separates claim-video consistency, assessed from internal evidence, from factual verdict determination, which additionally requires external corroboration. On top of it, a sufficiency-driven greedy search assembles evidence until a sufficiency threshold is met and outputs insufficient when the candidate pool is exhausted, rather than forcing a verdict. With Claude Sonnet 4, the method reaches a Macro-F1 of 0.510 using 4.5 evidence units on average (16% of the full evidence set), statistically indistinguishable from the full-evidence baseline (0.518 with 27.7 units), while significantly improving recognition of insufficient cases over the same search without abstention. The efficiency result replicates with GPT-5.5 and holds only partially with an open-weight Qwen2.5-72B verifier. Ablations show that external evidence is indispensable for factual determination, while internal video evidence grounds the verdict in claim-video consistency. These findings suggest that evidence-efficient verification is achievable, and that explicit abstention is needed when evidence is genuinely inadequate.
Comments33 pages, 2 figures, 20 tables