Maverick:通过矩阵-向量乘法委托实现实用化的私有且可验证的LLM推理
Maverick: Private and Verifiable LLM Inference Made Practical via Matrix-Vector Multiplication Delegation
浏览论文内容
中文总结 AI 辅助
Maverick提出基于矩阵-向量乘法委托的协议,实现信息论可验证且隐私保护的LLM推理,在Qwen3-4B上相比本地推理最高可达157倍加速。
中文摘要 AI 辅助
开源大型语言模型(LLM)在透明度方面日益与闭源模型竞争,同时能够在不将用户输入暴露给服务提供商的情况下运行推理。然而,在本地运行大规模模型需要大量的计算资源。在实践中,用户可能仍会求助于第三方提供商,从而引发隐私和正确性问题。现有的解决方案通常会给服务器带来巨大的开销或引入额外的信任假设。在本文中,我们提出了Maverick,一种基于矩阵-向量乘法委托协议的新型私有且可验证的LLM推理方法,矩阵-向量乘法是LLM中的主导操作。其核心在于,据我们所知,Maverick提供了第一个信息论上可靠的矩阵-向量乘法委托验证协议,具有透明的预处理、高效的(批量)验证以及几乎为零的服务器开销。我们将此验证原语与基于LPN的伪随机掩码相结合,以提供输入隐私。我们实现了矩阵-向量委托原语,并利用它构建了Maverick的端到端原型,我们通过在Qwen3-4B上测量每秒令牌数(tokens per second)的吞吐量来评估该原型。我们评估了具有1-8个线程的客户端配置。在单个客户端线程和最多使用128个线程的CPU服务器的情况下,当隐私掩码在线生成时,Maverick相对于本地推理的吞吐量提升高达17倍;当掩码预计算时,提升高达45倍;当仅需要验证时,提升高达44倍。在四个客户端线程的情况下,相应的提升分别为13倍、18倍和17倍。当服务器计算不再是瓶颈时,模拟网络延迟的客户端微基准测试显示加速比分别为12倍-20倍、34倍-135倍和38倍-157倍。
英文摘要
Open-source large language models (LLMs) are increasingly competitive with closed-source models while offering transparency and the ability to run inference without exposing user inputs to a service provider. However, running large-scale models locally requires substantial computational resources. In practice, users may still resort to a third-party provider, giving rise to privacy and correctness concerns. Existing solutions that address these problems often impose substantial server overhead or introduce additional trust assumptions. In this paper, we present Maverick, a novel approach to private and verifiable LLM inference based on a protocol for delegating matrix-vector multiplication, a dominant operation in LLMs. At its core, Maverick provides, to our knowledge, the first information-theoretically sound verification protocol for matrix-vector multiplication delegation with transparent preprocessing, efficient (batch) verification, and virtually no server overhead. We combine this verification primitive with LPN-based pseudorandom masking to provide input privacy. We implement our matrix-vector delegation primitive and use it to build an end-to-end prototype of Maverick, which we evaluate on Qwen3-4B by measuring throughput in tokens per second. We evaluate client configurations with 1-8 threads. With one client thread and a CPU server using up to 128 threads, Maverick achieves throughput gains over local inference of up to 17x when privacy masks are generated online, 45x when they are precomputed, and 44x when only verification is required. With four client threads, the corresponding gains are 13x, 18x, and 17x. When server computation is no longer the bottleneck, client-side microbenchmarks with simulated network delay show speedups of 12x-20x, 34x-135x, and 38x-157x.
发表机构
- Yale University(耶鲁大学)
- IC3
机构由 AI 辅助整理,请以论文原文为准。