arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

深入探讨统一视频保护在图像到视频和基于微调的定制方面的时间挑战

Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization

Yuxin Huang, Ziming Hong, Mingming Gong, Wanyu Wang, Jing Zhang, Tongliang Liu

arXiv 2607.13336首次发表:更新:

发表机构

The University of Sydney; University of Melbourne; City University of Hong Kong; Wuhan University; Mohamed bin Zayed University of Artificial Intelligence(悉尼大学; 墨尔本大学; 香港城市大学; 武汉大学; 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对基于扩散模型的视频定制引发的隐私等保护问题,提出TC-UAP方法,通过优化身份级多帧UAP并引入相关时间建模和损失,使其在时间上一致且鲁棒,在多种视频定制及攻击下实现强身份保护。

AI 中文摘要

近期基于扩散的视频生成模型实现了高质量个性化视频定制,引发隐私等保护问题。现有工作多聚焦图像保护,视频保护研究不足。保护视频面临图像级扰动难抗3D视频VAE时间压缩、单视频优化的扰动易受时间编辑且无法保护未见过的视频、时间不一致的扰动对时间攻击不鲁棒等挑战。为此提出TC-UAP,通过优化身份级多帧UAP,考虑视频VAE时间压缩引起的局部时间依赖性,引入内在时间建模和外在替代时间攻击损失,使扰动在时间上一致且对未见过的时间攻击鲁棒。实验表明,TC-UAP在基于参考和微调的视频定制下均实现最强身份保护,且在多种未见过的时间攻击下保持鲁棒。

英文摘要

Recent diffusion-based video generation models have enabled high-quality personalized video customization through both tuning-based pipelines, which fine-tune a video diffusion model, and reference-based pipelines such as image-to-video generation. However, these capabilities raise serious concerns about personal privacy, identity ownership and intellectual property protection. Existing anti-customization works focus on protecting images, while protection for videos against both reference- and tuning-based customization remains largely underexplored. Protecting videos in this setting raises three challenges: (i) Image-level perturbations, optimized frame by frame, cannot survive temporal compression by 3D video VAE. (ii) A video-level perturbation optimized on a single video is vulnerable to temporal editing and fails to protect unseen videos. (iii) Temporally inconsistent perturbations are not robust to temporal attacks. To address these challenges, we propose Temporally Consistent Universal Adversarial Perturbations (TC-UAP), the first protection method against both reference- and tuning-based video customization. TC-UAP optimizes an identity-level multi-frame UAP over sliding windows from multiple videos, accounting for local temporal dependencies induced by temporal compression in video VAE and enabling a single perturbation to protect unseen videos of varying lengths. Moreover, we introduce intrinsic temporal modeling and an extrinsic surrogate temporal-attack loss, which make the perturbation temporally consistent and robust to unseen temporal attacks. Empirically, quantitative and qualitative results show that TC-UAP achieves the strongest identity protection compared with existing methods under both reference- and tuning-based video customization, and remains robust under multiple unseen temporal attacks.

CommentsThis work provides a basis for the ECCV 2026 LifeGenIP Challenge on Unlearnable Videos against Diffusion-based Customization. Challenge page: https://lifegenip.cc/competition. Evaluation code: ECCV26_LifeGenIP_starting_kit" target="_blank" rel="noopener">https://github.com/tmllab/ECCV26_LifeGenIP_starting_kit. Project page: https://saythe17.github.io/TC-UAP/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑