arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18526cs.CR

本地隐私的幻象:消费级LLM服务系统中的机密性边界失效

The Illusion of Local Privacy: Confidentiality Boundary Failures in Consumer LLM Serving Systems

Youssef Hamdi Zafan Ibrahim, Muhammad Ikram, Mohammed Khalaf Salama

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示本地LLM推理并非天然私密,通过LLAnalyzer框架发现模型加载、内存、包装器及服务接口四类机密性失效,并报告授权缺陷与时序攻击,呼吁明确隐私保证。

中文摘要 AI 辅助

在本地运行大型语言模型(LLM)通常被认为比云托管推理更私密,因为用户提示词保留在设备上。我们探究仅将推理保持在本地是否足以保证这些提示词的机密性。我们的结果表明并非如此:提示词的机密性还取决于周围的 serving 软件在推理之前、期间和之后如何处理提示词数据。我们考察了消费级本地LLM服务系统中提示词机密性可能失效的四个边界:模型加载、运行时内存、包装器级持久化以及服务接口。为了研究这些边界,我们开发了 LLAnalyzer,一个测量框架,它分别测试每个边界,并将观察到的失效追溯到负责的软件组件。将 LLAnalyzer 应用于四个开放权重模型系列和两个消费级部署平台,我们发现不同边界的行为显著不同。在一场持续24小时的 AFL++ 活动中,执行超过1200万次,我们在所探索的状态空间内未观察到解析器崩溃或成功的畸形 GGUF 加载。运行时内存的情况则不同:我们在推理后恢复了提示词,因为分配器管理的内存中存在多个明文表示,而净化处理减少了这种残留但未完全消除。我们还发现,消费级包装器可以通过明文持久化延长提示词的生命周期。在服务边界,我们发现了 this http URL 中一个此前未被记录在案的授权缺陷,该缺陷允许一个已认证客户端恢复另一个租户保存的对话状态;该攻击在200/200次受控试验中均成功。此外,共享提示词前缀缓存暴露了一个远程时序预言机,在广域网条件下仍可区分。我们认为,本地LLM系统需要对提示词生命周期、持久化存储和租户隔离提供明确保证。

英文摘要

Running large language models (LLMs) locally is often considered more private than cloud-hosted inference because user prompts remain on the device. We ask whether keeping inference local is, by itself, sufficient to keep those prompts confidential. Our results show that it is not: prompt confidentiality also depends on how the surrounding serving software handles prompt data before, during, and after inference. We examine four boundaries at which prompt confidentiality can fail in consumer local-LLM serving systems: model loading, runtime memory, wrapper-level persistence, and the serving interface. To study these boundaries, we develop LLAnalyzer, a measurement framework that tests each boundary separately and traces observed failures to the responsible software component. Applying LLAnalyzer to four open-weight model families and two consumer deployment platforms, we find markedly different behaviour across boundaries. In a 24-hour AFL++ campaign with more than 12 million executions, we observe no parser crashes or successful malformed GGUF loads within the explored state space. Runtime memory tells a different story: we recover prompts after inference because multiple plaintext representations survive in allocator-managed memory, and sanitisation reduces this residue without eliminating it. We also find that consumer wrappers can extend prompt lifetime through plaintext persistence. At the serving boundary, we uncover a previously undocumented authorization flaw in llama.cpp that allows one authenticated client to restore another tenant's saved conversation state; the attack succeeds in 200/200 controlled trials. Separately, shared prompt-prefix caching exposes a remote timing oracle that remains distinguishable under WAN conditions. We argue that local LLM systems need explicit guarantees for prompt lifetime, persistent storage, and tenant isolation.

发表机构

  • Macquarie University(麦考瑞大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑