AI 中文总结
针对短视频假新闻检测的挑战,提出SRM-FND框架,通过自反思多模态推理及相关机制,在FakeSV和FakeTT数据集上实现优于强基线的检测性能。
AI 中文摘要
当前假新闻检测流程越来越多地利用大语言模型和视觉语言模型开展基于推理的分析,但仍存在若干未解决的挑战:在无真实思维链监督的情况下通过自反思提升推理质量、利用改进的推理能力助力下游模型微调、将单样本欺诈模式发现与多样本验证相连接。本文提出SRM-FND,一种用于短视频假新闻检测的自反思多模态推理框架。SRM-FND通过对比审议、迭代根本原因诊断和修正提示优化生成更高质量的推理;盲分析师、反结论推理器和自一致性仲裁器协同识别并保留判别性理由。该框架还融入双阶段、主题自适应视觉语言模型微调,以改进多模态 grounding 并实现轻量级主题专业化。对于不确定案例,它通过检索可信与可疑的共事件示例执行置信度驱动的多样本审查。在FakeSV和FakeTT数据集上的实验表明,SRM-FND优于强基线,产生更可靠、可解释的预测,并在跨数据集性能上实现显著提升。
英文摘要
Recent fake news detection pipelines increasingly leverage large language models and vision-language models for reasoning-based analysis. However, several challenges remain open: improving reasoning quality through self-reflection without ground-truth chain-of-thought supervision, using improved reasoning to benefit downstream model fine-tuning, and connecting single-sample fraudulent-pattern discovery with cross-sample verification. We propose SRM-FND, a self-reflective multimodal reasoning framework for short-video fake news detection. SRM-FND develops higher-quality reasoning through contrastive deliberation, iterative root-cause diagnosis, and corrective prompt refinement. A Blind Analyst, Counter-Conclusion Reasoner, and Self-Consistency Arbiter collaboratively identify and retain discriminative rationales. The framework also incorporates dual-phase, topic-adaptive vision-language model fine-tuning to improve multimodal grounding and enable lightweight topic specialization. For uncertain cases, it performs confidence-driven cross-sample review by retrieving credible and suspicious co-event examples. Experiments on FakeSV and FakeTT show that SRM-FND outperforms strong baselines, produces more reliable and interpretable predictions, and delivers noticeable improvements in cross-dataset performance.