基于条件潜在去噪的跨模态翻译用于视频深度伪造检测
Cross-modal Translation via Conditional Latent Denoising for Video Deepfake Detection
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对视频深度伪造检测中跨模态信息利用不足的问题,提出基于条件潜在去噪的跨模态翻译框架CTCLD,通过潜在空间双向去噪连接音视频分布,提升检测性能。
AI中文摘要:
视频深度伪造的日益增长的威胁需要多模态检测。除了作为真实性的独立指标外,音频和视觉信号具有内在的依赖性,这也为检测提供了关键标准。以往的方法往往忽视跨模态对应关系,阻碍了域间信息传递,并留下了未被探索的关键检测线索。为解决这一挑战,我们提出了一种名为基于条件潜在去噪的跨模态翻译(CTCLD)的框架,用于视频深度伪造检测。它在潜在空间中连接异构模态的不同分布,实现平滑的跨域信息传递以提高检测性能。我们首先通过分解音视频联合分布建立贝叶斯基础。随后,CTCLD通过双向潜在去噪在相互条件下翻译两种模态,有效捕捉操纵信号中的细微不一致性。实验结果表明,所提出的CTCLD实现了全面的域对齐,从而产生了一种具有竞争力的性能的鲁棒视频深度伪造检测方法。
英文摘要:
The growing threat of video deepfakes necessitates multimodal detection. Beyond serving as independent indicators of authenticity, audio and visual signals have intrinsic dependencies that also provide an essential criterion for detection. Previous methods often overlook the cross-modal correspondences, hindering information transfer between domains and leaving crucial detection cues unexplored. To address this challenge, we propose a framework called Cross-modal Translation via Conditional Latent Denoising (CTCLD) for video deepfake detection. It connects the distinct distributions of heterogeneous modalities in latent spaces, enabling smooth cross-domain information transfer to improve detection performance. We first establish a Bayesian foundation by decomposing the audio-visual joint distribution. Subsequently, CTCLD translates both modalities via bidirectional latent denoising conditioned on each other, effectively capturing subtle inconsistencies in the manipulated signals. Experimental results demonstrate that the proposed CTCLD enables comprehensive domain alignment, resulting in a robust video deepfake detection approach with competitive performance.