arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越向量隐藏:破解与缓解TEE卸载大语言模型中的共享方向权重混淆

Beyond Vector Hiding: Breaking and Mitigating Shared-Direction Weight Obfuscation in TEE-Offloaded Large Language Models

Menghui Zhang, Aoying Zheng, Guoxiao Liu, Zizhuang Deng, Jiejing Wen, Jincheng Zhuang, Ran Tao

arXiv 2608.26651首次发表:更新:

AI 中文总结

该研究针对TEE卸载LLM中的共享方向权重混淆,提出SpectralLeak、LatticeLeak攻击破解方案,并设计ButterflyCloak防御方案,实现高效破解与防御。

AI 中文摘要

受信任执行环境(TEE)对大语言模型(LLM)的屏蔽式分区,通过将混淆后的线性层卸载到不可信加速器,仅在TEE内保留少量修正来加速设备端推理。然而,早期轻量级混淆方案保留了权重向量方向,被ArrowMatch破解。为防御该攻击,ArrowCloak将同一隐藏方向的标量倍注入所有权重向量,实现轻量级可信修正。我们表明,这种复用在加速器可见的完整矩阵中留下了秩1关系。对于已发布的实值方案,我们提出SpectralLeak,其代理在12个任务设置中实现87.98%的平均准确率,而受害者模型的平均准确率为89.85%。在我们对ArrowCloak已发布模块化安全公式的防御友好型模Q实现中,模Q算术抑制了该频谱信号,但保留了模Q下的代数秩1关系。因此,我们提出LatticeLeak,利用由此产生的隐藏格。在我们的BERT-Base和GPT2-Base实验中,它精确重构了每个受保护的定点参数;在所有评估的架构中,重构后的模型保留了受害者级别的任务准确率,无需受害者查询、标签或微调。这些发现确定了共享秩1复用是我们攻击所利用的泄漏的根本原因。基于这一见解,我们设计了ButterflyCloak,一种带密钥的最大秩蝴蝶掩码,用不同的掩码行替换复用方向,同时保留快速可信修正。

英文摘要

Trusted Execution Environment (TEE)-shielded partitioning of Large Language Models (LLMs) accelerates on-device inference by offloading obfuscated linear layers to an untrusted accelerator while retaining only a small correction inside the TEE. However, earlier lightweight obfuscation schemes preserved weight-vector directions and were broken by ArrowMatch. To defend against this attack, ArrowCloak injects scalar multiples of the same hidden direction into all weight vectors, enabling lightweight trusted correction. We show that this reuse leaves a rank-one relation across the complete accelerator-visible matrix. For the released real-valued scheme, we propose SpectralLeak, which estimates and removes the shared component. Across 12 task settings, its surrogates achieve $87.98\%$ mean accuracy versus $89.85\%$ for the victims. In our defense-favorable mod-$Q$ realization of ArrowCloak's published modular security formulation, mod-$Q$ arithmetic suppresses this spectral signal but retains the algebraic rank-one relation modulo $Q$. We therefore propose LatticeLeak, which exploits the resulting hidden lattice. In our BERT-Base and GPT2-Base experiments, it reconstructs every protected fixed-point parameter exactly; across all evaluated architectures, the reconstructed models retain victim-level task accuracy without victim queries, labels, or fine-tuning. These findings identify shared rank-one reuse as the root cause of the leakage exploited by our attacks. Guided by this insight, we design ButterflyCloak, a keyed maximal-rank butterfly mask that replaces the reused direction with distinct mask rows while retaining fast trusted correction...

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑