arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25431cs.CR

GIFT:通过GPU信息流跟踪在LLM服务中强制实现用户数据隔离

Here is a GIFT: Enforcing User Data Isolation in LLM Serving via GPU Information Flow Tracking

Jiacheng Shi, Xunjie Wang, Cheng Tan, Jinyu Gu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出GIFT(GPU信息流跟踪系统),通过按用户加密、静态流分析等技术在LLM服务中实现用户数据隔离,扩展版GIFT-CC集成机密计算,在vLLM等框架上仅产生4-10.7%吞吐量开销。

中文摘要 AI 辅助

LLM服务框架在共享基础设施上处理大量用户数据,这些数据常包含敏感信息。确保共享同一服务框架(在CPU上)的用户与LLM算子(在GPU上)之间的隔离对隐私保护至关重要。本文提出GIFT,这是一个GPU信息流跟踪系统,可在LLM服务中以最小开销强制实现用户数据隔离。此外,GIFT的设计是非侵入式的,允许CPU端服务框架自由演进,它基于两个关键见解:第一,“加密即隔离”利用了CPU组件仅编排数据流而不操作内容的观察结果,因此按用户加密可在不修改服务逻辑的情况下提供隔离;第二,GPU内核表现出有限且可预测的信息流,支持静态流分析。GIFT为每个内核预计算信息流规则并采用解耦的流跟踪,避免了插桩或GPU停顿。此外,我们将GIFT扩展为GIFT-CC,它集成了机密计算以防范不可信的操作系统和 hypervisor(LLM服务提供商)。在vLLM和DistServe上实现后,GIFT和GIFT-CC以4-10.7%的吞吐量开销实现了用户数据隔离,同时保持相同的延迟水平。

英文摘要

LLM serving frameworks process large volumes of user data--often containing sensitive information--on shared infrastructure. Ensuring isolation between users who share the same serving framework (on CPUs) and LLM operators (on GPUs) is critical for privacy protection. This paper presents GIFT, a GPU Information Flow Tracking system that enforces user data isolation in LLM serving with minimal overhead. Moreover, the design of GIFT is non-intrusive and allows CPU-side serving frameworks to evolve freely. It rests on two key insights. First, encryption-as-isolation leverages the observation that CPU components only orchestrate data flow, not content manipulation; thus, per-user encryption can provide isolation without modifying serving logic. Second, GPU kernels exhibit limited and predictable information flows, enabling static flow analysis. GIFT precomputes information flow rules for each kernel and uses decoupled flow tracking, avoiding instrumentation or GPU stalls. Furthermore, we extend GIFT to GIFT-CC, which integrates confidential computing to protect against untrusted operating systems and hypervisors (LLM service providers). Implemented on vLLM and DistServe, GIFT and GIFT-CC enforce user data isolation with a 4-10.7% throughput overhead while maintaining the same latency level.

↑