arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.04498cs.CVcs.SD

UniSkip-Mamba:用于视听时间伪造定位的频率感知状态空间模型

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization

Cangjin Yu, Quan Zhang, Dan Jiang, Ke Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对视听时间伪造定位,通过频域分析发现伪造特征集中在低中频,提出融合多模态序列与新颖扫描机制的UniSkip-Mamba框架,实现性能提升和推理加速。

中文摘要 AI 辅助

随着人工智能生成内容的激增,复杂的多媒体操纵引发了对恶意应用的严重关注,使视听时间伪造定位成为紧迫的研究前沿。现有方法在基于Transformer的时间建模和通道级多模态融合方面取得了进展,但处理所有频率成分时存在过拟合等问题。通过系统的频域分析,我们发现伪造判别模式集中在低/中频范围,高频成分主要引入噪声。基于此,我们提出了UniSkip-Mamba,一种频率感知状态空间模型框架,它结合了统一多模态序列融合以保留跨模态相位关系,并通过新颖的组扫描合并机制实现频率感知正则化的跳过扫描曼巴块,自然地将学习偏向于判别性的低/中频模式,同时保持表示完整性。我们实现了当前最优性能:在LAV-DF上AP@0.95为63.4%(提高9.8%),在AV-Deepfake1M上mAP为63.58%(提高14.32%),推理速度快6倍。我们的频域分析从信号处理角度为跳过扫描为何能提高准确性和鲁棒性提供了理论依据。

英文摘要

With the proliferation of AI-generated content, sophisticated multimedia manipulation has raised critical concerns about malicious applications such as opinion manipulation and evidence fabrication, making Audio-Visual Temporal Forgery Localization (AV-TFL) an urgent research frontier. Existing TFL methods have progressed along two main paradigms: Transformer-based temporal modeling and channel-wise multimodal fusion. While these approaches capture temporal dependencies and cross-modal correlations, they process all frequency components indiscriminately, leading to overfitting on high-frequency noise and limited robustness under real-world data degradation. Through systematic frequency domain analysis, we find that forgery-discriminative patterns concentrate in the low/mid-frequency range (normalized frequency 0-0.15), while high-frequency components primarily introduce noise, removing them even improves detection performance by +1.4%. Based on this phenomenon, we propose UniSkip-Mamba, a frequency-aware State Space Model framework that incorporates Unified Multimodal Sequence Fusion to preserve cross-modal phase relationships, and Skip-Scanning Mamba Blocks that implement frequency-aware regularization through a novel Group-Scan-Merge mechanism, naturally biasing learning toward discriminative low/mid-frequency patterns (0-0.15) while maintaining representational completeness. We achieve state-of-the-art (SOTA) performance: 63.4% AP@0.95 on LAV-DF (+9.8% improvement) and 63.58% mAP on AV-Deepfake1M (+14.32% improvement), with 6x faster inference. Our frequency-domain analysis provides theoretical justification from a signal processing perspective for why skip-scanning inherently improves both accuracy and robustness.

发表机构

  • Soochow University(苏州大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑