ROSETTA:通过混合CKKS/TFHE评估实现高效且准确的隐私保护LLM解码
ROSETTA: Efficient and Accurate Privacy-Preserving LLM Decoding via Hybrid CKKS/TFHE Evaluation
- Peking University(北京大学)
- Xi’an Jiaotong University(西安交通大学)
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对全同态加密下LLM解码中非线性运算开销大的问题,提出混合CKKS/TFHE框架ROSETTA,通过自适应分段查找表和方案感知算子选择,实现高达4.8倍Softmax加速和1.5-2.1倍端到端加速。
AI中文摘要:
生成式大语言模型(LLMs)在代码生成和问答等许多现实任务中已取得最先进的性能。这些模型主要依赖自回归解码策略,顺序生成输出令牌。然而,它们的广泛部署引发了严重的隐私问题,促使基于全同态加密(FHE)的私有推理框架应运而生。现有FHE框架的一个主要限制是评估非线性运算的效率低下,这些运算产生大量开销并主导解码阶段。在本文中,我们提出ROSETTA,一种混合CKKS/TFHE框架,以克服这一限制。我们首先观察到解码阶段的非线性运算表现出异构的工作负载模式,这可以通过混合方法有效处理。然后我们通过两个关键贡献实现这一点:1)基于TFHE的自适应分段查找表协议,能够高效且准确地评估非线性运算;2)一种方案感知的算子选择框架,自动将每个非线性算子分配给CKKS或TFHE,以最小化端到端解码延迟。我们证明ROSETTA在Softmax上实现了高达4.8倍的加速,在端到端延迟上实现了1.5至2.1倍的加速,优于最先进的框架CacheMir。
英文摘要:
Generative large language models (LLMs) have achieved state-of-the-art performance on many real-world tasks such as code generation and question answering. These models predominantly rely on an autoregressive decoding strategy that generates output tokens sequentially. However, their pervasive deployment raises serious privacy concerns, motivating private inference frameworks based on fully homomorphic encryption (FHE). A major limitation of existing FHE frameworks is their inefficiency in evaluating nonlinear operations, which incur substantial overhead and dominate the decode stage. In this paper, we propose ROSETTA, a hybrid CKKS/TFHE framework that overcomes this limitation. We first observe that nonlinear operations in the decode stage exhibit heterogeneous workload patterns, which can be handled effectively via a hybrid approach. We then realize this with two key contributions: 1) an adaptive segmented lookup-table protocol based on TFHE that enables efficient and accurate evaluation of nonlinear operations; and 2) a scheme-aware operator-selection framework that automatically assigns each nonlinear operator to CKKS or TFHE to minimize end-to-end decoding latency. We demonstrate that ROSETTA achieves up to $4.8\times$ Softmax speedup and $1.5$--$2.1\times$ end-to-end speedup over the SOTA framework CacheMir.