发表机构
ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MOSAIC提出新颖矩阵乘法掩码协议,通过误差缩放机制解决噪声累积问题,可高效安全外包大型Transformer推理,实现大规模机密AI计算。
AI 中文摘要
我们解决了在以下场景中,将AI计算从可信但计算能力弱的客户端安全且高效地外包给不可信但强大的服务器的挑战:客户端持有输入和模型,而服务器不得学习任何相关信息。我们提出了MOSAIC,其核心是一种新颖的矩阵乘法掩码协议,该协议可扩展到比现有工作大得多的矩阵,从而支持大型Transformer推理等现代工作负载的安全外包。通过在乘法结果中引入少量噪声并因此放宽正确性要求,MOSAIC实现了最优的渐近客户端开销,且具体运行时比现有工作快几个数量级。其安全性归约为决策性LWE和LPN假设。由于这种噪声会在Transformer的多层中累积,关键技术挑战是限制误差增长;MOSAIC通过基于随机Hadamard旋转的误差缩放机制解决了这一问题。在700亿参数的大型Transformer模型上,MOSAIC的困惑度与流行的量化方法相当,甚至在HumanEval数据集上与全精度BF16推理的表现匹配。最后,我们提出了端到端实现,展示了类似MOSAIC的思路如何为现代数据中心中的大规模机密AI提供可行路径。非机密推理已通过按阶段(预填充/解码)、层和时间分布,利用类RDMA网络在节点间移动激活值、缓存的KV值和权重,以最大化异构硬件的利用率。MOSAIC通过保持可信计算基(TCB)较小,并将大部分AI计算外包给不可信加速器,实现了机密计算的扩展。
英文摘要
We address the challenge of securely and efficiently outsourcing AI computations from a trusted but computationally weak client to an untrusted but powerful server, in the setting where the client holds both the input and the model, and the server must learn neither. We present MOSAIC, whose core is a novel matrix-multiplication masking protocol that scales to far larger matrices than prior work, enabling the safe outsourcing of modern workloads such as large transformer inference. By introducing small amounts of noise to the multiplication result and thereby relaxing correctness, MOSAIC achieves optimal asymptotic client overhead and concrete runtimes orders of magnitude faster than prior work. Its security reduces to the decisional LWE and LPN assumptions. Because this noise accumulates across the many layers of a transformer, a key technical challenge is bounding error growth; MOSAIC addresses this with an error-scaling mechanism based on random Hadamard rotations. On large 70B transformer models, MOSAIC's perplexity is comparable to popular quantization approaches and even matches full-precision BF16 inference on HumanEval. Finally, we present an end-to-end implementation showing how ideas like MOSAIC can promise a path towards large-scale confidential AI in modern data centers. Non-confidential inference is already distributed across phase (prefill/decode), layer, and time to maximize utilization of heterogeneous hardware, using RDMA-like networking to move activations, cached KV values, and weights across nodes. MOSAIC enables scaling of confidential compute by keeping the trusted computing base (TCB) small and outsourcing the bulk of the AI computation to untrusted accelerators.