发表机构
Yangzhou University; Auburn University(扬州大学; 奥本大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Genda生成式弹幕框架与DM-FEND检测模型,解决弹幕延迟导致的假新闻检测实时性问题,在FakeSV、FakeTT基准上性能优于现有最优方法,为多模态假新闻检测提供鲁棒方案。
AI 中文摘要
现代多媒体平台上,观众通过弹幕(又称子弹评论)进行的社交互动既会引发观点冲突也能形成共识,提供可用于假新闻检测的细粒度判别性社交信号。但现实场景中弹幕固有的累积延迟违背了假新闻检测的实时性需求,导致相关研究不足。为解决该问题,本文提出一种新型时间感知生成式弹幕框架Genda,模拟用户互动过程,该框架包含两部分:(1)弹幕触发器,用于预测用户反应的时机与强度;(2)弹幕生成器,用于合成对应的语义与情感表达,从而共同构建时间对齐、类人的伪弹幕流。为使生成的弹幕可用于识别假新闻视频,本文进一步设计了弹幕引导时间多模态假新闻检测模型DM-FEND,该模型实现视频、音频、文本与弹幕间的细粒度多模态交互,增强动态模态对齐与语义噪声抑制。实验结果表明,在中文基准FakeSV和英文基准FakeTT上,DM-FEND均持续优于现有最优基线方法;进一步的 ablation 实验验证了时间弹幕建模对提升鲁棒性和判别能力的关键作用。最后,本研究通过弥合新闻与用户行为间的时间不一致性,为现代社交互动场景下的多模态假新闻检测提供了一种有效且鲁棒的解决方案。
英文摘要
The social interactions among crowds via \textit{Danmaku} (a.k.a., bullet comments) on modern multimedia platforms can facilitate both viewpoint conflicts and consensus, providing fine-grained discriminative social signals that can benefit fake news detection. However, the inherent accumulation latency of \textit{Danmaku} in real-world scenarios violates the real-time necessity of fake news detection, making the studies of \textit{Danmaku}-related fake news detection underexplored. To break this violation, we simulate this temporal-aware user interactive process by proposing a novel temporal \textbf{Gen}erative \textbf{da}nmaku framework, called \textbf{Genda}, which consists of: (1) a \textit{Danmaku} Trigger for predicting the timing and intensity of user reactions; and (2) a \textit{Danmaku} Generator for synthesizing corresponding semantic and emotional expressions, thereby mutually constructing a temporally aligned and human-like pseudo \textit{Danmaku} streams. To make the generated \textit{Danmaku} useful for identifying fake news videos, we further design a \textit{Danmaku}-guided Temporal Multimodal fake news detection model - \textbf{DM-FEND}, which enables fine-grained multimodal interactions among video, audio, text, and \textit{Danmaku}, enhancing dynamic modalities alignment and semantic noise inhibition. The experimental results demonstrate that \emph{DM-FEND} consistently outperforms state-of-the-art baselines across both Chinese (FakeSV) and English (FakeTT) benchmarks. Further ablations validate the crucial role of temporal \textit{Danmaku} modeling in enhancing robustness and discriminative capability. Finally, this study offers a bright and robust solution for multimodal fake news detection in modern social interactive fashions by bridging the temporal inconsistency between news and user behaviors.