发表机构
Peking University; The University of Western Australia; Shandong University(北京大学; 西澳大学; 山东大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
JITterFlip是首个针对基于GPU的LLM推理主机侧JIT服务控制平面的位翻转攻击,可生成乱码或正确输出,能绕过现有防御,通过Rowhammer攻击实现跨边界传播,在四个LLM上取得高PPL或延迟放大效果。
AI 中文摘要
大语言模型(LLMs)广泛通过云托管推理服务部署,其中即时编译(Just-in-Time, JIT)用于降低框架和GPU启动的重复开销。JIT服务引入了主机侧控制平面,用于选择编译产物并协调其在GPU上的执行。与此同时,共享云环境催生了针对LLM/DNN推理的大量位翻转攻击(Bit-Flip Attacks, BFAs)。现有大多数位翻转攻击针对模型参数或权重,需要特定模型的知识;少数工作通过故障可执行代码减少了这种依赖,但仍会破坏直接实现模型计算的代码,将攻击效果限制在推理耗竭。本文提出JITterFlip,首个针对基于GPU的LLM推理的主机侧JIT服务控制平面的位翻转攻击。通过故障CPU驻留的服务决策而非模型计算,JITterFlip既能生成乱码输出,也能实现正确输出的海绵攻击。为在大型JIT编译器栈中识别可利用目标,JITterFlip开发了决策引导的故障脆弱代码分析。在四个文本和多模态LLM工作负载中,识别出的脆弱代码故障表现出跨模型可迁移性,产生的乱码输出困惑度(PPL)比值为15.45倍至2.48×10^6倍,正确输出的海绵攻击延迟放大倍数为2.03倍至181.90倍。JITterFlip还能绕过近期针对LLM的位翻转攻击防御,同时保留两种攻击效果。最后,我们展示了针对四个LLM的端到端Rowhammer攻击:CPU驻留分支代码中的单个位翻转跨CPU-GPU边界传播,在不直接访问GPU内存的情况下破坏GPU执行的推理,达到高达7.23×10^6倍的PPL放大或124.97倍的延迟放大,同时保留完全相同的生成输出。
英文摘要
LLMs are widely deployed through cloud-hosted inference services, where Just-in-Time (JIT) compilation is used to reduce recurring framework and GPU-launch overhead. JIT serving introduces a host-side control plane that selects compiled artifacts and orchestrates their execution on the GPU. Meanwhile, the shared cloud setting has motivated a growing body of bit-flip attacks (BFAs) against LLM/DNN inference. Most existing BFAs target model parameters or weights and require model-specific knowledge. A smaller body of work reduces this dependency by faulting executable code, yet still corrupts code that directly implements model computation, limiting their attack effect to inference depletion. We present JITterFlip, the first BFA targeting the host-side JIT serving control plane of GPU-based LLM inference. By faulting CPU-resident serving decisions rather than model computation, JITterFlip enables both gibberish output generation and a correct-output sponge attack. To identify exploitable targets in a large JIT compiler stack, JITterFlip develops a decision-guided fault-vulnerable code analysis. Across four text and multimodal LLM workloads, the identified vulnerable code faults exhibit cross-model transferability, produce gibberish outputs with PPL ratios of $15.45\times$ to $2.48{\times}10^{6}\times$, and demonstrate correct-output sponge attacks with latency amplification of $2.03\times$ to $181.90\times$. JITterFlip also bypasses recent BFA defenses for LLMs while retaining both attack effects. Last, we demonstrate end-to-end Rowhammer attacks across four LLMs: a single bit flip in CPU-resident branch code propagates across the CPU-GPU boundary to disrupt GPU-executed inference without direct access to GPU memory, reaching up to $7.23{\times}10^{6}\times$ PPL amplification or $124.97\times$ latency amplification while preserving the exact generated output.