发表机构
Tsinghua University; Zhongguancun Laboratory; Institute for Advanced Study, Tsinghua University; School of Cryptographic Science and Engineering, Shandong University; State Key Laboratory of Cryptography and Digital Economy Security, Tsinghua University; Shandong Institute of Blockchain, Shandong(清华大学; 中关村实验室; 清华大学高等研究院; 山东大学密码科学与工程学院; 清华大学密码与数字经济安全国家重点实验室; 山东省区块链研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种猜测-确定框架,通过识别维度猜测留下的零后缀和相等模式,联合恢复ReLU全连接网络的架构与参数,首次实现无需已知架构假设的密码分析提取攻击。
AI 中文摘要
密码分析提取攻击仅通过黑盒访问神经网络的原始输出即可恢复其参数。然而,所有现有攻击都依赖一个基本假设:攻击者知道网络架构。例如,对于基于ReLU激活的全连接网络,网络深度和每个隐藏层的维度是已知的。在本文中,我们研究是否可以移除这一假设。我们专注于ReLU全连接网络,并提出一个猜测-确定框架,该框架联合恢复架构和参数。我们方法的核心是一个简单但强大的观察:维度猜测会在参数恢复过程中留下架构敏感的痕迹。我们识别出两种这样的痕迹:(i)签名恢复产生的合并权重向量中的零后缀,其长度揭示了多余猜测的数量;(ii)基于原像的符号恢复中的相等模式,该模式仅在维度猜测正确时出现。这两个信号产生了两种互补的恢复路径。我们进一步提出了两个标准来识别倒数第二层,这对于终止猜测过程是必要的。我们在广泛的ReLU网络上实现了端到端攻击,包括扩张和非扩张架构。据我们所知,这是第一个移除已知网络架构假设的密码分析提取攻击。
英文摘要
Cryptanalytic model extraction aims to reconstruct a functionally equivalent model through black-box interactions with the victim model. Under the fundamental assumption that the network architecture is completely known, existing attacks achieve the goal by recovering the model parameters. In this paper, we explore whether this assumption can be removed practically. Focusing on ReLU fully connected networks, which are widely studied in this field, we propose a guess-and-determine framework that jointly recovers the network architecture (including network depth and hidden-layer dimensions) and the model parameters. This framework is based on a simple yet effective high-level idea: after designing a parameter recovery attack under the known-architecture assumption, we can analyze the architecture-sensitive traces observed during parameter recovery to recover the network architecture. We identify two such traces in differential extraction attacks: (i) a \emph{zero suffix} in the merged weight vectors produced by signature recovery, whose length reveals the hidden layer dimension; and (ii) an \emph{equality pattern} in the preimage-based sign recovery, which occurs only under the true hidden layer dimension. These two signals give rise to two routes for network architecture recovery. For the second-to-last layer, we further propose two methods, one for identifying it, and one for recovering its dimension. Practical end-to-end attacks are implemented on a wide range of ReLU neural networks, including both expansive and non-expansive networks. To the best of our knowledge, this is the first time the feasibility of achieving functionally equivalent extraction on deep neural networks, after removing the known-architecture assumption, has been demonstrated in practice.