arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00742cs.CVcs.MM

注意裂缝:人工智能生成视频检测中的跨尺度耦合不匹配

Mind the Rift: Cross-Scale Coupling Mismatch for AI-Generated Video Detection

Siyu Li, Jin Yang, Weiheng Liang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对AI生成视频检测难题,提出RIFT框架,利用跨尺度耦合不匹配作为取证信号,在两个基准数据集上实现高F1分数,且具有编码器无关性。

中文摘要 AI 辅助

随着AI视频生成器达到电影级真实感,可靠的检测对于维护数字信任变得至关重要。我们将跨尺度耦合不匹配识别为一种新的取证信号,其中尺度指的是抽象水平(语义动态与像素级残差):在自然视频中,宏观层面的时间动态与微观层面的残差模式通过统一的成像物理流程内在耦合,而AI生成器的训练目标并未明确保留这种联合分布,因此会系统性地破坏这种耦合。检测此类不匹配颇具挑战,因为它需要分别提取两个尺度的信息,同时量化它们的跨尺度关系。我们提出RIFT(Representation Inconsistency Forensics on Trajectories,即轨迹表示不一致取证),这是一种正交取证框架,通过三个相互关联的组件解决该问题:宏观流通过微分几何和学习到的流形轨迹上的持续同伦构建预期时间演化的动态基线;微观流通过隐写分析滤波和时间建模充当敏感的取证探针;耦合发散模块测量两个流之间的条件依赖性。格拉姆-施密特正交性保证了该测量的信息论有效性。在两个基准数据集(VidProM,含12万条视频、7种生成器;GenVidBench,含6.8万条视频、4种生成器)上的实验表明,RIFT分别达到99.33%和99.72%的F1分数,在留一生成器评估中实现97.87%的未见生成器检测率,同时表现出编码器无关性:从ViT-S/14(2200万参数)扩展到ViT-L/14(3亿参数)时,F1分数变化小于0.1%,切换至不同编码器家族(DINOv1)仅使F1分数降低0.73个百分点。代码可在指定的URL获取。

英文摘要

As AI video generators achieve cinematic realism, reliable detection becomes essential for safeguarding digital trust. We identify cross-scale coupling mismatch as a new forensic signal, where scale refers to the level of abstraction (semantic dynamics vs. pixel-level residuals): in natural videos, macro-level temporal dynamics and micro-level residual patterns are intrinsically coupled by the unified imaging physics pipeline, whereas AI generators, whose training objectives do not explicitly preserve this joint distribution, systematically violate this coupling. Detecting such mismatch is challenging because it requires independently extracting information at both scales while simultaneously quantifying their cross-scale relationship. We propose RIFT (Representation Inconsistency Forensics on Trajectories), an orthogonal forensic framework that addresses this through three interlocking components: a macro stream that builds a dynamic baseline of expected temporal evolution via differential geometry and persistent homology on learned manifold trajectories, a micro stream that acts as a sensitive forensic probe via steganalytic filtering and temporal modeling, and a coupling divergence module that measures the conditional dependency between the two streams. Gram-Schmidt orthogonality guarantees the information-theoretic validity of this measurement. Experiments on two benchmarks (VidProM, 120K videos, 7 generators; GenVidBench, 68K videos, 4 generators) demonstrate that RIFT achieves 99.33% and 99.72% F1-score respectively, with 97.87% unseen-generator detection rate in leave-one-out evaluation, while exhibiting encoder agnosticism: scaling from ViT-S/14 (22M) to ViT-L/14 (300M) changes F1 by less than 0.1%, and switching to a different encoder family (DINOv1) reduces F1 by only 0.73 pp. Code is available at https://github.com/Litsay/RIFT

补充信息

↑