ARMOR++:用于对深度伪造检测器进行可转移攻击的多域原语集的智能编排
ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
- Department of Informatics, Aristotle University of Thessaloniki(塞萨洛尼基亚里士多德大学信息学系)
- Infocomm Technology Cluster, Singapore Institute of Technology(新加坡科技学院信息通信技术集群)
- Department of Applied Physics and Applied Mathematics, Columbia University(哥伦比亚大学应用物理与应用数学系)
- Division of Natural and Applied Sciences, Duke Kunshan University(昆山杜克大学自然科学与应用科学部)
- Department of Electrical and Computer Engineering, University of Toronto(多伦多大学电气与计算机工程系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对深度伪造检测器在黑盒对抗转移下可靠性下降问题,提出ARMOR++多智能体框架,利用视觉与语言模型提供语义先验并编排原语选择等,整合多种原语有效针对异构偏差,实验表明其性能优于现有基线,凸显当前检测器可靠性差距及智能编排的有效性。
AI中文摘要:
深度伪造检测器的可靠性在黑盒对抗转移下经常下降,因为这些模型通常依赖脆弱的、与架构相关的取证线索。现有转移攻击缺乏语义感知,在严格的无查询约束下难以保持有效性。本文介绍了ARMOR++,一个用于高转移性深度伪造逃避的强大多智能体框架。该框架利用Qwen2.5-VL视觉语言模型提供空间语义先验,Qwen3大语言模型编排原语选择、自适应超参数重新参数化和熵正则化扰动混合。通过整合五种互补原语,ARMOR++有效针对异构归纳偏差。在AADD-2025基准上的严格评估表明,ARMOR++在低质量和高质量图像模式下均显著优于现有智能体和非智能体基线。统计分析证实,与最先进的智能体基线相比,盲目标攻击成功率有大幅提高,在针对非智能体基准和强大防御配置下也有进一步性能优势。这些发现凸显了当前深度伪造检测器部署中存在的显著可靠性差距,并证明了智能编排识别潜在漏洞的有效性。
英文摘要:
The reliability of deepfake detectors frequently degrades under black-box adversarial transfer, as these models often rely on fragile, architecture-dependent forensic cues. Existing transfer attacks often lack semantic awareness and struggle to maintain effectiveness under strict no-query constraints, particularly when perturbations are transferred from convolutional surrogates to transformer-based targets. To address these limitations, this paper introduces ARMOR++, a robust multi-agent framework designed for high-transferability deepfake evasion. The framework leverages the Qwen2.5-VL Vision-Language Model (VLM) to supply spatial semantic priors, while the Qwen3 Large Language Model (LLM) orchestrates primitive selection, adaptive hyperparameter reparameterization, and entropy-regularized perturbation mixing. By integrating five complementary primitives, spanning dense optimization, saliency-based methods, spatial transformations, frequency-domain perturbations, and block-structured modifications, ARMOR++ effectively targets heterogeneous inductive biases. Rigorous evaluation on the AADD-2025 benchmark demonstrates that ARMOR++ significantly outperforms existing agentic and non-agentic baselines across both low- and high-quality image regimes. Statistical analysis confirms a substantial gain in blind-target Attack Success Rate (ASR) over the state-of-the-art agentic baseline, with further performance advantages evidenced against non-agentic benchmarks and under robust defensive configurations. These findings highlight a significant residual reliability gap in current deepfake detector deployments and demonstrate the efficacy of agentic orchestration in identifying latent vulnerabilities.