arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深度伪造新闻检测:一种集成LipNet、DeepSpeech和ResNET的多模态框架用于增强视听分析

Deepfake News Detection: A Multimodal Framework Integrating LipNet, DeepSpeech and ResNET for Enhanced Audio-Visual Analysis

Ameena Khan, Muhammad Ahsan Aziz, Muhammad Junaid Asif, Naeem Akhter, Rana Fayyaz Ahmad

arXiv 2607.20579首次发表:更新:

AI 中文总结

针对深度伪造新闻威胁数字新闻媒体真实性的问题,提出集成LipNet、DeepSpeech和ResNET的多模态框架,从音频和视觉线索提取特征,经多种模型分类,实验表明该方法准确率达94%,优于基线,具鲁棒性和实际潜力。

AI 中文摘要

深度伪造新闻是指通过操纵面部表情或语音来欺骗观众的人工智能生成(或人工智能操纵)的多媒体内容。生成式人工智能的快速发展使得高度逼真的假视频和克隆语音的合成变得广泛,对数字新闻媒体的真实性构成严重威胁。本文提出了一个多模态框架,通过联合利用音频和视觉线索来辨别视频内容的真实性,以应对检测深度伪造视频的挑战。我们提出的框架包括从唇部动作、音频内容和视频帧中提取特征。唇部动作和语音内容分别使用LipNet和DeepSpeech2模型进行编码,面部特征通过BlazeFace提取并用ResNet18表示。提取的特征向量被连接成一个整体的视频表示,并使用包括随机森林(RF)、多层感知器(MLP)和长短期记忆(LSTM)网络在内的机器学习和深度学习模型进行分类。在FakeAVCeleb数据集上进行的大量实验表明,所提出的方法在使用增强音频特征时达到了94%的准确率,优于当前最先进的多模态集成基线。结果证实了所提出框架在深度伪造新闻检测方面的鲁棒性和实际潜力。

英文摘要

Deepfake news refers to AI-generated (or AI ma-nipulated) multimedia content intentionally generated to deceive audiences by manipulating the facial expressions, or speech while maintaining the realistic appearance. The rapid progress of generative AI has made the synthesis of highly realistic fake videos and cloned voices widely accessible, posing a serious threat to the authenticity of digital news media. This paper presents a multi-modal framework that discerns the authenticity of video content by jointly exploiting audio and visual cues, thereby addressing the challenge of detecting the deepfake videos. We proposed a framework that involves features extraction from lip movements, audio content and video frames. Lip movements and speech content are encoded using the LipNet and DeepSpeech2 models, while facial features are extracted by leveraging the use of BlazeFace and represented with ResNet18. The extracted feature vectors are concatenated into a holistic video representation and classified with an ensemble of machine learning and deep learning models, including Random Forest (RF), Multi-layer Perceptron (MLP) and Long Short-Term Memory (LSTM) networks. Exten-sive experiments performed on the FakeAVCeleb dataset shows that the proposed approach attains an accuracy of 94% using augmented audio features, outperforming a state-of-the-art multi-modal ensemble baseline. The results confirm the robustness and practical potential of the proposed framework for deepfake news detection.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑