arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

端到端硬标签密码分析模型提取:利用高效符号恢复

End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery

Akira Ito, Takayuki Miura, Yosuke Todo

arXiv 2609.21941首次发表:更新:

发表机构

Tohoku University; NTT Social Informatics Laboratories(东北大学; NTT社会信息学实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于全新原理的符号恢复算法,无需专用查询,实现高效符号恢复,从而在完全黑盒设置下对训练好的深度ReLU MLP进行端到端硬标签模型提取,在MNIST和Fashion-MNIST上达到98%以上标签一致性。

AI 中文摘要

深度神经网络(DNN)的重要性已被广泛认可,通过训练获得的参数被视为宝贵资产。近年来,在IACR会议上,仅利用对DNN的预言机查询来提取这些参数的攻击已被积极研究。硬标签设置是模型提取中最具挑战性的设置,其中对手只能观察到最终输出标签,如“狗”或“猫”。在Eurocrypt 2025上,Carlini等人提出了基于ReLU的MLP的多项式时间硬标签提取方法。然而,该攻击过程的一个步骤,即符号恢复,需要大量查询和大量计算。在黑盒设置中实现此步骤仍然困难。因此,在训练好的深度ReLU MLP上进行完全黑盒的端到端演示仍然是一个挑战。在本文中,我们提出了一种新的符号恢复算法,其原理与现有方法完全不同。我们的方法不需要专门用于符号恢复的查询。在我们的实验中,它比现有方法实现了更高的符号恢复准确率。因此,即使对于训练好的模型,它也能实现高效的符号恢复。借助我们的符号恢复算法,硬标签模型提取的所有步骤都可以在黑盒设置中实现。通过结合这些实现,我们展示了从在MNIST和Fashion-MNIST上训练的模型(宽度为16,具有4或6个隐藏层)进行端到端模型提取,实现了超过98%的标签一致性。

英文摘要

The importance of deep neural networks (DNNs) is widely recognized, and the parameters obtained through training are regarded as valuable assets. Recently, attacks that extract these parameters using only oracle queries to a DNN have been actively studied at IACR conferences. The hard-label setting is the most challenging setting for model extraction, where an adversary can observe only the final output label, such as "dog" or "cat." At Eurocrypt 2025, Carlini et al. proposed polynomial-time hard-label extraction of ReLU-based MLPs. However, one step of this attack process, i.e., sign recovery, requires a large number of queries and substantial computation. Implementing this step in a black-box setting remains difficult. Consequently, a fully black-box end-to-end demonstration on trained deep ReLU MLPs has remained a challenge. In this paper, we propose a new sign-recovery algorithm based on a completely different principle from the existing method. Our method requires no dedicated queries for sign recovery. In our experiments, it achieves higher sign-recovery accuracy than the existing method. Consequently, it enables efficient sign recovery even for trained models. With our sign-recovery algorithm, all steps of hard-label model extraction can be implemented in a black-box setting. By combining these implementations, we demonstrate end-to-end model extraction from models trained on MNIST and Fashion-MNIST, with width 16 and 4 or 6 hidden layers, achieving over 98% label agreement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑