arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38246cs.CRcs.AI

ModalFidelity:在预算约束下为深度伪造检测路由模态

ModalFidelity: Routing Modalities for Deepfake Detection on a Budget

Oguzhan Baser, Kaan Kale, Sriram Vishwanath, Sandeep Chinchali

首次发表
浏览论文内容

中文总结 AI 辅助

针对深度伪造隐藏于视频未知小片段的问题,提出轻量级路由器ModalFidelity,在硬计算预算下预检窗口并选择值得读取的流,以15.9倍更少计算量达到更高准确率,并保留神谕96%以上性能。

中文摘要 AI 辅助

深度伪造不再需要伪造整个视频。能够读取转录文本的生成器现在只修改视频中含义转变的那几秒钟,因此伪造内容隐藏在整个视频中一个未知的小片段里。然而,检测器仍然会读取音频和图像流的每一个一秒窗口,将近乎全部的计算量花费在没有被篡改的部分。我们观察到,决定看哪里远比实际去看要便宜。我们提出了ModalFidelity,一个轻量级路由器,它在任何取证检测器运行之前预览每个窗口,并决定在无法超越的硬计算预算下,哪个流值得读取。在AV-Deepfake1M上,最多读取五分之一的窗口,其准确率高于在检测器之后进行门控的方法,同时计算量减少15.9倍,并且保留了知道每个伪造位置的神谕(oracle)超过96%的准确率。

英文摘要

Deepfakes no longer need to fake a whole video. Generators that read the transcript now alter only the few seconds in which a video's meaning turns, so a forgery hides in a small, unknown fraction of the video. Yet detectors still read every one-second window of both the audio and image streams, spending nearly all of their compute where nothing was altered. We observe that deciding where to look is far cheaper than looking. We present ModalFidelity, a lightweight router that previews each window and decides, before any forensic detector runs, which stream is worth reading, under a hard compute budget it can never exceed. On AV-Deepfake1M, reading at most a fifth of the windows, it is more accurate than gating after the detectors at 15.9x less compute, and retains over 96% of the accuracy of an oracle that knows where every forgery lies.

发表机构

  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑