arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过模型重编程防御图神经网络的模型提取攻击

Defending against Model Extraction for GNNs with Model Reprogramming

Yan Wen, Zhenyi Wang, Heng Huang

arXiv 2608.11495首次发表:更新:

发表机构

University of Maryland, College Park; University of Central Florida(马里兰大学帕克分校; 中佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对图神经网络的模型提取攻击,提出主动防御框架GraphRP,通过结构感知门控机制构建动态结构防火墙,可显著降低攻击有效性且保留良性效用。

AI 中文摘要

图神经网络(GNNs)是机器学习即服务(MLaaS)中高风险应用的核心,但它们的黑盒部署使其面临模型提取(ME)攻击,攻击者通过查询API窃取知识产权。现有防御存在严重的“欧几里得偏差”:将基于图像的策略(如随机噪声)迁移到图上,却忽略节点间复杂的拓扑依赖,常导致效用严重下降;水印等被动方法也无法实时防止盗窃。为弥合这一差距,我们提出GraphRP(图重编程保护),一种将模型重编程用于安全的主动防御框架。与静态扰动不同,GraphRP引入由可学习拓扑原型驱动的结构感知门控机制,构建动态“结构防火墙”,选择性调节模型决策边界:保留训练流形上良性查询的保真度,同时最大化对抗查询在扰动方向上的费舍尔信息。在标准假设(有界损失、最优攻击者、局部二阶近似)下,我们证明攻击者估计误差的下界随重编程噪声的结构敏感性增加而增大。对硬标签和软标签ME攻击的大量实验表明,GraphRP可显著降低攻击有效性,同时保留良性效用。

英文摘要

Graph Neural Networks (GNNs) serve as the backbone for high-stakes applications in Machine-Learning-as-a-Service (MLaaS). Still, their black-box deployment exposes them to Model Extraction (ME) attacks, in which adversaries steal intellectual property by querying APIs. Existing defenses suffer from a critical ''Euclidean bias'': they transfer image-based strategies (e.g., random noise) to graphs, ignoring the complex topological dependencies between nodes, which often results in severe utility degradation. Passive methods like watermarking also fail to prevent theft in real time. To bridge this gap, we propose GraphRP (Graph Reprogramming Protection), a proactive defense framework that repurposes Model Reprogramming for security. Unlike static perturbations, GraphRP introduces a Structure-Aware Gating Mechanism driven by learnable topological prototypes. This creates a dynamic ''structural firewall'' that selectively modulates the model's decision boundary: it preserves fidelity for benign queries residing on the training manifold, while maximizing the Fisher Information along the perturbation direction for adversarial queries. Under standard assumptions (bounded loss, optimal attacker, and local second-order approximation), we prove a lower bound on the attacker's estimation error that increases with the structural sensitivity of the reprogramming noise. Extensive experiments on both hard-label and soft-label ME attacks demonstrate that GraphRP significantly degrades attack effectiveness while preserving benign utility.

CommentsAccepted by KDD 2026

DOI:10.1145/3770855.3817983

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑