arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10117cs.CR

基于混淆的端侧大语言模型保护的安全边界理解

Understanding the Security Boundary of Obfuscation-based On-Device LLM Protection

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Hanyi Zhou, Chenyang Li, Yuanzhe Pang, Ke Xu, Mingwei Xu, Zhuotao Liu

中文总结 AI 辅助

本文针对基于TEE的端侧LLM保护中启发式混淆方法的局限,形式化混淆原语以统一先前TSLP方法,刻画其安全边界,提出Collapse攻击揭示共同漏洞,并引入新原语扩展安全边界。

中文摘要 AI 辅助

可信执行环境(TEE)为保护端侧大语言模型(LLM)的知识产权提供了一种有前景的机制。为了克服TEE固有的计算瓶颈,现有的TEE保护型LLM分区(TSLP)方法对计算密集型层应用高效的混淆方案,将其卸载到外部GPU,而仅在TEE内保留轻量级操作。尽管基于TSLP的方法日益增多,但这些防御机制在很大程度上仍是启发式的。因此,一些方法被证明易受某些专门设计的、旨在利用其特定架构实现的对抗性攻击的影响。为克服这些启发式设计的局限性,本文探讨了一个基础性研究问题:我们能否建立通用原语来统一具有代表性的先前方法,刻画其组合的安全边界,并系统地扩展它们?为此,我们形式化了一组混淆原语,定义为满足特定代数性质的线性计算二元组。我们证明,本文研究的代表性高效TSLP框架的矩阵级权重变换可以表示为这些原语的组合;因此,这些原语组合的规范形式(记为O_prior)刻画了该原语族的结构边界。随后,我们通过一种新颖的原语引导攻击方法Collapse揭示了O_prior的脆弱性,证明了多个发表在顶级会议上的著名TSLP方法(如ArrowCloak(Security'25)、TSQP(S&P'25)和LoRO(NeurIPS'25))存在共同漏洞。最后,我们引入了两种新颖的混淆原语,并将其与现有构造相结合以形成O_ext,从而扩展了这一安全边界。

英文摘要

Trusted Execution Environments (TEEs) offer a promising mechanism for safeguarding the intellectual property of on-device Large Language Models (LLMs). To overcome the inherent computational bottlenecks of TEEs, existing TEE-Shielded LLM Partition (TSLP) methods apply efficient obfuscation schemes to computationally intensive layers, offloading them to external GPUs while retaining only lightweight operations within the TEE. Although a growing body of TSLP-based approaches has emerged, these defense mechanisms remain largely heuristic. Consequently, some methods are proven vulnerable to certain specialized adversarial attacks designed to exploit their specific architectural implementations. To overcome the limitations of these heuristic designs, this paper addresses a fundamental research question: can we establish common primitives to unify representative prior methodologies, characterize the security boundary of their compositions, and systematically extend them? To this end, we formalize a set of obfuscation primitives, defined as dual-tuples of linear computations satisfying specific algebraic properties. We demonstrate that the matrix-level weight transformations of several representative efficient TSLP frameworks can be expressed as compositions of these primitives; consequently, the canonical form of these primitive compositions, denoted as \priorboundary, defines the security boundary of this primitive family. We then expose the vulnerabilities of \priorboundary through a novel primitive-guided attack methodology, \sysattack, demonstrating a shared vulnerability in several prominent TSLP methods published in top-tier venues, such as ArrowCloak (Security'25), TSQP (S\&P'25), and LoRO (NeurIPS'25). Finally, we introduce two novel obfuscation primitives and integrate them with existing constructs to formulate \sysdefense, extending the prior security boundary \priorboundary.

补充信息

↑