发表机构
Hong Kong Baptist University; Beijing Normal University(香港浸会大学; 北京师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对纯合成T2V假新闻视频带来的新威胁,构建了首个纯合成假新闻视频数据集,提出R-T2V框架,在10种主流基准上取得最优检测性能。
AI 中文摘要
近期,文本生成视频(T2V)模型可从零合成假新闻视频,使威胁超越了由现有素材拼接而成的廉价伪造视频。这类新闻视频能与虚构叙事高度匹配,给现有检测器造成模态对齐陷阱。现有数据集缺乏纯合成假新闻视频。尽管直接用假新闻视频描述提示T2V模型可产生完全对齐的样本,但这会将假新闻视频检测(FNVD)简化为单模态捷径,引发语义-视觉退化。为解决该问题,我们将T2V-FNVD定义为包含真实、廉价伪造、纯合成伪造三类标签的新型三元分类任务,并构建首个纯合成假新闻视频数据集(PS-FNVD)。PS-FNVD包含两类样本:一是具备对齐欺骗性的虚构事件(1型),二是具有虚假视觉来源的真实事件(2型),可防止模型利用单模态捷径。此外,我们提出推理引导的T2V-FNVD(R-T2V)框架,该框架通过条件理由生成和监督微调训练,将高层语义逻辑与低层物理生成痕迹相结合,以预测三元真实性标签。对10种主流基准方法开展的大量实验显示,R-T2V实现了最优性能,准确率较次优基准高出12.20个百分点,宏F1值高出8.46个百分点。
英文摘要
Recent text-to-video (T2V) generation models enable fake news videos to be synthesized from scratch, shifting the threat beyond cheap fakes assembled from existing footage. Such news videos can closely match fabricated narratives, creating a modality alignment trap for existing detectors. Existing datasets lack pure synthesis fake news videos. Although directly prompting T2V models with descriptions of fake news videos can yield perfectly aligned samples, it reduces the fake news video detection (FNVD) to unimodal shortcuts and causes semantic-visual degeneration. To counter this, we formulate T2V-FNVD as a novel ternary classification task with three labels (real, cheap fake, and pure synthesis fake) and construct the first pure synthesis fake news video dataset (PS-FNVD). PS-FNVD includes fabricated events with aligned deception (Type 1) and true events with false visual provenance (Type 2), preventing models from exploiting unimodal shortcuts. Furthermore, we propose the Reasoning-guided T2V-FNVD (R-T2V) framework. Trained through conditioned rationale generation and supervised fine-tuning, R-T2V integrates high-level semantic logic with low-level physical generative traces to predict the ternary veracity label. Extensive experiments across 10 prevailing baselines show that R-T2V achieves the state-of-the-art performance, outperforming the second-best baseline by 12.20 percentage points in accuracy and 8.46 percentage points in macro $F_1$.
CommentsAccepted at ACM Multimedia (ACM MM), 2026