ZK-LLM 下的比特:评估面向可验证私有 LLM 推理的零知识友好量化
Bits Under ZK-LLM: Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference
- Siebel School of Computing and Data Science(西贝尔计算与数据科学学院)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文首次系统研究 LLM 的零知识友好量化,发现激活精度比权重更敏感,RMSNorm 查找是瓶颈,且降低位宽不保证证明成本成比例下降,提出算子感知精度选择。
AI中文摘要:
零知识证明正成为一种有前景的方法,用于实现私有、可验证的 LLM 治理与审计,其中监管者、用户和审计者需要验证关于训练数据使用或 LLM 推理时行为的声明,而模型提供者必须保护专有模型参数。然而,尽管对 ZK-LLM 的兴趣日益增长,对 ZK 友好量化的理解仍然有限。这一差距很重要,因为在 ZK 设置中,量化直接塑造了 ZK 推理的算术结构、约束复杂性和证明成本。ZK 协议在有限域上运行,其成本在很大程度上取决于算术运算、非线性操作和查找约束的数量与类型。因此,理解 ZK 友好量化对于使 ZK-LLM 实用化至关重要。在这项工作中,我们首次对 LLM 的 ZK 友好量化进行了系统性研究。我们首先形式化了 ZK 友好量化的定义,捕捉了 ZK 证明生成所需的属性。然后,我们在权重、激活和非线性查找表精度的广泛设计空间中,评估了九个语言模型,包括 Qwen2.5-14B 和混合专家模型 Qwen3-30B-A3B。我们的结果表明,激活精度比权重精度敏感得多,而非线性查找近似可能成为效用下降的主要来源。此外,我们识别出 RMSNorm 逆平方根查找是多个大型模型中的反复出现的瓶颈,并通过仅在瓶颈处选择性地提高精度恢复了接近基线的效用。最后,我们表明降低位宽或查找表大小并不一定能带来成比例的端到端证明节省,这表明传统的低位量化启发式方法不能直接转化为 ZK 证明效率,并激励了算子感知的精度选择。
英文摘要:
Zero-knowledge proofs are emerging as a promising approach for enabling private, verifiable LLM governance and auditing, where regulators, users, and auditors need to verify claims about training-data usage or LLM inference-time behavior, while model providers must protect proprietary model parameters. However, despite the growing interest in ZK-LLMs, the understanding of ZK-friendly quantization remains limited. This gap matters because in the ZK setting, quantization directly shapes the arithmetic structure, constraint complexity, and proving cost of ZK inference. ZK protocols operate over finite fields and incur costs that depend heavily on the number and type of arithmetic operations, nonlinearities, and lookup constraints. Understanding ZK-friendly quantization is therefore essential for making ZK-LLMs practical. In this work, we present the first systematic study of ZK-friendly quantization for LLMs. We first formalize the definition of ZK-friendly quantization, capturing the properties required for ZK proof generation. We then evaluate nine language models, including Qwen2.5-14B and the mixture-of-experts model Qwen3-30B-A3B, across a broad design space of weight, activation, and nonlinear lookup table precision. Our results show that activation precision is substantially more sensitive than weight precision, while nonlinear lookup approximations can become the dominant source of utility degradation. Also, we identify RMSNorm inverse-square-root lookups as a recurring bottleneck in several large models and recover near-baseline utility by selectively increasing precision only at the bottleneck. Finally, we show that reducing bit-width or lookup-table size does not necessarily yield proportional end-to-end proving savings, showing that conventional low-bit quantization heuristics do not directly translate to ZK proving efficiency and motivating operator-aware precision selection.