LLMscope:通过光学探测从边缘AI芯片中提取大语言模型(LLM)资产
LLMscope: Extracting LLM Assets from Edge AI Chips via Optical Probing
- Worcester Polytechnic Institute(伍斯特理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出LLMscope方法,通过激光电压成像攻击基于FPGA的LLM加速器,可从边缘AI芯片提取LLM资产,建立了资产恢复方法及成像工作量与资产维度的下界
AI中文摘要:
将大语言模型(LLM)推理迁移至边缘AI加速器会带来新的物理漏洞。执行期间,模型参数与中间推理状态会反复加载到芯片并在其上处理,易受物理侧信道攻击。本研究通过部署激光电压成像,证明可在推理过程中从局部存储器和计算子电路中提取LLM资产,即嵌入向量、注意力机制、量化多层感知机(MLP)权重、激活值及其他推理状态。为验证上述结论,我们对基于FPGA的LLM加速器实施攻击;由于此类加速器在地址、块、模块及层间复用相同的缓冲器与计算子电路,读取资产值等价于在推理期间探测不同存储器。我们展示了目标值的完整恢复;同时建立了即使部分权重或位未被读取仍可恢复资产值的方法。我们进一步推导了关联成像工作量与资产维度的下界,证明即便直接恢复也会随目标资产规模呈线性缩放
英文摘要:
The move of LLM inference to edge AI accelerators introduces new physical vulnerabilities. During execution, model parameters and intermediate inference states are repeatedly loaded into and processed on the chip, making them suscep- tible to physical side-channel attacks. In this work, by deploying laser voltage imaging, we show that one can extract LLM assets during inference, namely embeddings, attention, and quantized MLP weights, activations, and other inference states, from localized memories and compute subcircuits. To validate our claims, we perform an attack on an FPGA-based LLM accelerator. Since such accelerators reuse the same buffers and compute subcircuits across addresses, tiles, modules, and layers, reading asset values comes down to probing different memories during inference. We demonstrate full recovery of the targeted values; however, we also establish a methodology to recover asset values even if some weights or bits remain unread. We further derive lower bounds that relate imaging effort to asset dimensions and show that even direct recovery scales linearly with the size of the targeted asset