保障野外大语言模型安全:边缘环境下的隐私与安全挑战
Securing LLMs in the Wild: Privacy and Security Challenges at the Edge
浏览论文内容
中文总结 AI 辅助
研究边缘大语言模型安全隐私挑战,围绕内存墙等架构约束引入分类法,得出统一约束模型,提出安全操作效率得分,给出决策程序和缓解措施,为联合评估安全、隐私和效率提供框架,保障边缘原生智能系统。
中文摘要 AI 辅助
大语言模型正迅速从研究环境走向野外,部署在企业基础设施、个人设备和边缘平台上。云部署虽提供可扩展计算,但数据主权、合规性、延迟和第三方依赖等问题促使组织采用边缘和本地大语言模型。这种转变带来新的安全和隐私挑战,如有限的计算和内存促使激进优化,可能引入漏洞并重塑威胁格局,即安全-效率悖论。本文研究了压缩如何降低安全一致性、分区推理如何引发重建攻击以及持续本地适应如何导致隐私泄露和模型漂移。通过围绕内存墙、二次墙和计算墙这三个架构约束引入以部署为中心的分类法,得出统一约束模型,量化不安全优化何时不可避免,并提出安全操作效率得分(SOES),还给出实用决策程序和针对性缓解措施,为联合评估安全、隐私和效率提供协同设计框架,为保障边缘原生智能系统奠定基础。
英文摘要
Large Language Models (LLMs) are rapidly moving from research settings into the wild, deployed on enterprise infrastructure, personal devices, and edge platforms. While cloud deployments offer scalable compute, concerns over data sovereignty, compliance, latency, and third-party dependence are driving organizations toward edge and on-premise LLMs. This shift introduces new security and privacy challenges: limited compute and memory force aggressive optimizations, including quantization, pruning, model partitioning, and parameter-efficient adaptation, each of which can introduce vulnerabilities and reshape the threat landscape. We describe this tension as the Security-Efficiency Paradox, mechanisms that improve efficiency may weaken robustness, expose new attack surfaces, or increase privacy risks. We examine how compression can degrade safety alignment, how partitioned inference enables reconstruction attacks, and how continuous local adaptation may cause privacy leakage and model drift. To analyze these risks, we introduce a deployment-centric taxonomy organized around three architectural constraints: the Memory Wall, the Quadratic Wall, and the Compute Wall. We derive a unified constraint model that quantifies when unsafe optimizations become unavoidable, linking each wall to specific attack surfaces. Building on this model, we propose the Secure Operational Efficiency Score (SOES), a holistic metric balancing task accuracy, jailbreak resistance, and privacy against energy, memory, and latency, enabling practitioners to configure edge LLMs under real-world hardware limits. We further present a practical decision procedure and targeted mitigations for each optimization-induced vulnerability. Together, these contributions provide a co-designed framework for jointly evaluating security, privacy, and efficiency, laying a foundation for securing edge-native intelligent systems.
发表机构
- Department of Electrical and Computer Engineering, University of South Florida(电子与计算机工程系,佛罗里达州立大学)
- Department of Mathematics, Embry-Riddle Aeronautical University(数学系,埃默里-瑞德航空大学)
机构由 AI 辅助整理,请以论文原文为准。