arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于人工智能篡改视频检测的集成深度学习方法

Ensemble Deep Learning Approaches for AI-Altered Video Detection

Laiba Khan, Hung-Mao Wu, Wei Lin, Frank Bi, Yousef Abdelhadi, Joshua Jung

arXiv 2607.06872首次发表:更新:

发表机构

Department of Mathematical and Computational Sciences, University of Toronto Mississauga; Department of Mathematical and Computational Sciences, University of Toronto; University of Toronto(多伦多大学密西沙加分校数学与计算科学系; 多伦多大学数学与计算科学系; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对人工智能篡改视频检测难题,开发多模态深度伪造检测系统,结合音频视觉分析,用模型集成及均值平均、堆叠等策略组合分数,提升检测鲁棒性与性能,凸显对未见篡改泛化的挑战,平均准确率约70%。

AI 中文摘要

人工智能的日益普及导致人工智能生成的视频迅速增加,使得区分真实内容和被篡改内容变得更加困难。许多现有检测方法依赖单一模型,难以在不同类型的深度伪造上实现泛化。在这项工作中,我们开发了一种多模态深度伪造检测系统,使用模型集成结合音频和视觉分析。该系统包括用于基于音频检测的AASIST,以及用于分析视频帧视觉特征的EfficientNet、XceptionNet和MesoNet。该流程将视频作为输入,分离音频,并使用MTCNN提取人脸帧。每个模型生成一个分数表示输入为伪造的可能性。然后使用包括均值平均和堆叠在内的集成策略组合这些分数。均值融合提供了一个简单稳定的基线,而堆叠使用训练好的元模型来学习如何更有效地组合预测。结果表明,虽然单个模型在其训练的数据集上表现良好,但在更多样化的数据集上测试时性能会下降。集成方法通过组合多个模型的预测有助于提高整体鲁棒性,在不同类型的深度伪造上实现更一致的性能。这表明同时使用音频和视觉信息是一种更可靠的深度伪造检测方法。我们的结果突出了对未见篡改的泛化作为核心开放挑战,平均准确率约为70%。

英文摘要

The increasing accessibility of artificial intelligence has led to a rapid rise in AI-generated videos, making it more difficult to distinguish between real and manipulated content. Many existing detection methods rely on a single model and often struggle to generalize across different types of deepfakes. In this work, we developed a multimodal deepfake detection system that combines both audio and visual analysis using an ensemble of models. The system includes AASIST for audio-based detection, and EfficientNet, XceptionNet, and MesoNet for analyzing visual features in video frames. The pipeline takes a video as input, separates the audio, and extracts face frames using MTCNN. Each model produces a score indicating the likelihood of the input being fake. These scores are then combined using ensemble strategies, including mean averaging and stacking. Mean fusion provides a simple and stable baseline, while stacking uses a trained meta-model to learn how to combine predictions more effectively. Results show that while individual models perform well on the datasets they were trained on, their performance drops when tested on more diverse datasets. The ensemble approach helps improve overall robustness by combining predictions from multiple models, leading to more consistent performance across different types of deepfakes. This suggests that using both audio and visual information together is a more reliable approach for deepfake detection. Our results highlight generalization to unseen manipulations as the central open challenge, with average accuracy around 70%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑