arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于良性目的的对抗性攻击:视觉内容全生命周期主动保护综述

Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle

Jiaming Zhang, Boyang Chen, Zherui Li, Fuyao Zhang, Xinyu Yan, Hong Xi Tae, Wenwen He, Xuan Wang, Siqi Guo, Junhao Dong, Kun Wang, Hanxun Huang, Yige Li, Xingjun Ma, Yang Cao, Lingjuan Lyu, Wei Yang Bryan Lim

arXiv 2608.04314首次发表:更新:

发表机构

College of Computing and Data Science, Nanyang Technological University; School of Computing and Information Systems, University of Melbourne; School of Computing and Information Systems, Singapore Management University; Institute of Trustworthy Embodied AI, Fudan University; Department of Computer Science, School of Computing, Institute of Science Tokyo; Sony AI(南洋理工大学计算与数据科学学院; 墨尔本大学计算与信息系统学院; 新加坡管理大学计算与信息系统学院; 复旦大学可信具身人工智能研究院; 东京科学大学计算学院计算机科学系; 索尼人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述研究了视觉内容全生命周期中用于主动保护的“良性对抗性攻击”范式,涵盖五类方法并评估其特性,同时指出当前保护措施的不足与未来研究方向。

AI 中文摘要

一旦视觉内容进入AI流程,其所有者通常对内容的使用方式几乎没有技术控制权。法律和监管补救措施可解决滥用问题,但许多技术干预措施必须在更早阶段应用,即内容发布或访问时。本综述研究围绕这一干预点发展起来的保护范式,我们称之为“用于良性目的的对抗性攻击”。长期以来被研究作为攻击学习模型的扰动和结构化信号,现在被数据所有者、创作者、平台或审计员用来扰乱未授权自动化或支持后续问责。五个研究领域大致独立地实现了这种反转,每个领域针对视觉资产生命周期的不同阶段:共享时用于防止未授权识别的隐私过滤器、针对未授权训练的不可学习示例、针对恶意编辑或模仿的生成式安全措施、用于访问控制以对抗自动化代理的对抗性验证码,以及用于传播后归因的来源机制。尽管这些方法在不同场合开发且成功标准不兼容,但许多方法利用了人类感知、语义解释与机器推理之间的持续差距,表明随着视觉流程向多模态模型和自主代理发展,该范式仍具相关性。为使各方法的主张具有可比性,我们沿可迁移性、适应性和部署就绪性的共享轴对这五类方法进行评估。在整个生命周期中,我们发现大多数保护措施仍主要针对静态或弱适应性对手进行验证,而受控基准之外的证据仍很稀缺。最后,我们整合跨阶段对策以及针对稳健、可组合且可部署的所有者侧保护的开放问题。

英文摘要

Once visual content enters an AI pipeline, its owner often retains little technical control over how it is used. Legal and regulatory remedies can address misuse, but many technical interventions must be applied earlier, when content is released or accessed. This survey examines the protective paradigm that has grown around this intervention point, which we call \emph{adversarial attacks for good}. Perturbations and structured signals long studied as attacks on learned models are instead applied by data owners, creators, platforms, or auditors to disrupt unauthorized automation or support later accountability. Five research communities have arrived at this inversion largely independently, each addressing a different stage of a visual asset's lifecycle: privacy filters against unwanted recognition at sharing time, unlearnable examples against unauthorized training, generative safeguards against malicious editing or imitation, adversarial CAPTCHAs for access control against automated agents, and provenance mechanisms for post-circulation attribution. Although developed in separate venues with incompatible success criteria, many of these methods exploit persistent gaps between human perception, semantic interpretation, and machine inference, suggesting that the paradigm remains relevant as visual pipelines evolve toward multimodal models and autonomous agents. To make their claims comparable, we evaluate all five families along shared axes of transferability, adaptability, and deployment readiness. Across the lifecycle, we find that most protections are still validated mainly against static or weakly adaptive adversaries, while evidence beyond controlled benchmarks remains scarce. We close by consolidating cross-stage countermeasures and open problems for robust, composable, and deployable owner-side protection.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑