平滑Transformer前馈网络的曲率密码分析
Curvature Cryptanalysis of Smooth Transformer Feed-Forward Networks
- National Institute of Standards and Technology(美国国家标准与技术研究院)
- Michigan Technological University(密歇根理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对带GELU/SiLU激活的Transformer前馈网络,利用黑盒二阶观测的曲率泄漏,仅用少量查询即可恢复其内部结构,生成高保真替代模型。
AI中文摘要:
我们证明,平滑的两层前馈网络(FFN)在FFN分支处的选定输入原始输出预言机下,存在额外的结构模型提取通道;研究Transformer的FFN分支,其使用GELU或SiLU激活函数,攻击者仅能访问选定输入的原始输出,无法获取参数、梯度或内部激活值;利用二阶泄漏通道,其中投影输入的海森矩阵(Hessians)形成由FFN输入权重诱导的相同隐藏对称秩1因子的不同混合。我们将所得海森矩阵收集形式化为部分对称分解,以建立局部可识别性和稳定性的条件,从而利用向量输出模板复用将结构查询成本降低16倍。在独立训练的CIFAR-10视觉Transformer上,仅需16个投影海森矩阵(对应8193次黑盒查询)即可恢复隐藏的FFN方向,平均绝对余弦对齐度超过0.94,其中95.1%的GELU方向和91.9%的SiLU方向对齐度超过0.90。在独立训练的模型、重复提取运行及所有Transformer块中,恢复效果均保持较高水平。恢复的结构也支持功能提取:将恢复的方向固定,仅拟合剩余FFN参数即可生成高保真替代模型,其Top-1准确率超过93%,测试准确率与GELU和SiLU目标的差距分别为0.90%和0.62%。在固定攻击配置下,输出舍入和高斯噪声会大幅降低恢复效果,但调整有限差分步长可将平均对齐度恢复至0.9603和0.9398。这是从黑盒二阶观测到隐藏FFN结构恢复及功能替换的端到端路径,在所述预言机模型下,平滑FFN的曲率暴露了仅通过行为保真度无法揭示的内部参数几何结构。
英文摘要:
We show that smooth two-layer feed-forward networks (FFNs) expose an additional structural model extraction channel under a chosen-input raw-output oracle at the FFN branch; consider transformer FFN branches with GELU or SiLU activations under chosen-input raw-output access, without access to parameters, gradients, or internal activations; exploit a second-order leakage channel in which projected input Hessians form different mixtures of the same hidden symmetric rank-one factors induced by the FFN input weights. We formalize resulting Hessian collection as a partially symmetric decomposition to establish conditions for local identifiability and stability to exploit vector-output stencil reuse to reduce the structural query cost by a factor of 16. On independently trained CIFAR-10 vision transformers, only 16 projected Hessians, corresponding to 8193 black-box queries, recover the hidden FFN directions with average absolute cosine alignment above 0.91, with 86.7.1 % of GELU and 91.1 % of SiLU directions exceeding 0.90 alignment. Recovery remains high across independently trained models, repeated extraction runs, and all transformer blocks. The recovered structure supports functional extraction too. Keeping the recovered directions fixed and fitting only the remaining FFN parameters yields high-fidelity substitutes with more than 92 % top-1 agreement, while test accuracy remains within 1.24% and 0.57% of the GELU and SiLU targets. Output rounding and Gaussian noise substantially reduce recovery under a fixed attack configuration, but adapting the finite-difference step restores average alignment to 0.9188 and 0.9081. This is an end-to-end path from black-box second-order observations to hidden FFN-structure recovery and functional replacement. Under the stated oracle model, smooth FFN curvature exposes internal parameter geometry that behavioral fidelity alone cannot reveal.