引导流:通过梯度引导流匹配实现人脸识别模型的逆攻击
Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching
浏览论文内容
中文总结 AI 辅助
本文提出两阶段白盒模型逆攻击方法SFMI,通过梯度引导流匹配将逆攻击转化为轨迹引导任务,在人脸数据集上实现了优于现有方法的攻击性能。
中文摘要 AI 辅助
模型逆攻击(MIAs)旨在从人脸识别模型中重建目标身份的代表性训练样本,暴露了严重的安全漏洞。现有方法通常依赖间接引导或高度随机的引导,难以将生成轨迹稳定优化至目标人脸图像。本文提出了引导流模型逆攻击(SFMI),这是一种新颖的两阶段白盒模型逆攻击方法,将逆攻击重新表述为轨迹引导任务。具体而言,第一步是学习通用流匹配先验,预训练一个通用的无条件Flow Matching模型,将人脸流形编码为鲁棒先验;第二步是使用渐进式引导调度器(PGS)进行攻击,在采样过程中注入随时间变化的目标特定梯度。PGS通过反向传播目标模型以获取中间生成状态的梯度,逐步将自适应引导信号注入向量场,该过程有效引导当前生成流从随机噪声走向目标类别的高密度区域。在使用CelebA数据集的身份不相交交叉评估设置下,SFMI在ArcFace目标上达到了0.9248的准确率(ACC)、22.61的Fréchet inception距离(FID)和0.3874的学习感知图像块相似度(LPIPS)。对多个目标模型的大量实验表明,在评估的白盒协议下,SFMI在攻击成功率和视觉保真度方面达到了具有竞争力的SOTA性能。
英文摘要
Model Inversion Attacks (MIAs) aim to reconstruct representative training samples of target identities from face recognition models, exposing critical security vulnerabilities. Existing methods typically rely on indirect guidance or highly stochastic guidance, making it difficult to stably optimize generation trajectories toward target facial images. In this paper, we propose Steering Flow Model Inversion (SFMI), a novel two-stage white-box model inversion method that reformulates inversion as a trajectory-steering task. Specifically, Step I, Learning a Generic Flow Matching Prior, pre-trains a generic unconditional Flow Matching model to encode the manifold of human faces as a robust prior. Step II, Attacking with Progressive Guidance Scheduler (PGS), injects time-dependent target-specific gradients during sampling. By backpropagating through the target model to obtain gradients from intermediate generated states, PGS progressively injects adaptive guidance signals into the vector field. This process effectively steers the current generative flow from random noise toward the high-density regions of the target class. Under an identity-disjoint cross-evaluation setting using the CelebA dataset, SFMI achieves an ACC of 0.9248, an FID of 22.61, and an LPIPS of 0.3874 on the ArcFace target. Extensive experiments on multiple target models demonstrate that SFMI achieves competitive state-of-the-art performance in attack success and visual fidelity under the evaluated white-box protocol.
发表机构
- School of Cyberspace Science, Harbin Institute of Technology(哈尔滨工业大学网络空间安全学院)
机构由 AI 辅助整理,请以论文原文为准。