发表机构
Nanchang University; Beihang University(南昌大学; 北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种融合空间与频率域特征的视频伪造检测模型,基于ResNet-LSTM框架结合CBAM和DCT,在多个基准数据集上验证了优越性能,为复杂条件下的视频真实性分析提供了可靠方案。
AI 中文摘要
随着信息技术的进步,数字内容已在新闻广播、娱乐、商业和法医调查等多个领域得到广泛应用。然而,先进多媒体编辑工具的普及显著增加了视频和图像伪造的风险,在社会和个人层面对内容真实性提出了严重关切。为满足对稳健且准确的检测方法日益增长的需求,本研究提出了一种新颖的视频伪造检测模型,该模型融合了空间域和频率域特征。该模型基于ResNet-LSTM框架构建,并通过卷积块注意力模块(CBAM)增强空间特征提取,同时引入离散余弦变换(DCT)以捕获频率域信息。研究在多个主流基准数据集上进行了全面实验,涵盖了广泛的伪造场景。结果表明,所提模型在区分真实与篡改视频方面表现出优越性能。额外的消融研究和对比研究证实了架构中每个组件的贡献,为模型能力提供了更深入的见解。总体而言,研究结果支持所提方法作为在复杂条件下增强视频真实性分析可靠性的有前景的解决方案。
英文摘要
As information technology advances, digital content has become widely adopted across diverse fields such as news broadcasting, entertainment, commerce, and forensic investiga?tion. However, the availability of sophisticated multimedia editing tools has significantly increased the risk of video and image forgery, raising serious concerns about content authenticity at both societal and individual levels.To address the growing need for robust and accurate detection methods, this study proposes a novel video forgery detection model that integrates both spatial and frequency-domain features. The model is built on a ResNet-LSTM framework enhanced by a Convolutional Block Attention Module (CBAM) for spatial feature extraction, and further incorporates Discrete Cosine Transform (DCT) to capture frequency domain information. Comprehensive experiments were conducted on several mainstream benchmark datasets, encompassing a wide range of forgery scenarios. The results demonstrate that the proposed model achieves superior performance in distinguishing between authentic and manipulated videos. Additional ablation and comparative studies confirm the contribution of each component in the architecture, offering deeper insight into the models capacity. Overall, the findings support the proposed approach as a promising solution for enhancing the reliability of video authenticity analysis under complex conditions.