arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01730cs.CRcs.AI

HEAT:通过近似-权重协同适配实现更快的全同态推理

HEAT: Faster Fully Homomorphic Inference via Approximations-Weights Co-Adaptation

Alessandro Zirilli, Davide Marincione, Evgenios M. Kornaropoulos, Giuseppe Ateniese, Emanuele Rodolà

首次发表
浏览论文内容

中文总结 AI 辅助

HEAT 是一种同态加密感知训练方法,通过让非线性激活的迭代次数与模型权重协同适配,在 GPT-2 加密解码任务上显著降低了迭代次数、引导次数与端到端延迟,同时提升了解码一致性。

中文摘要 AI 辅助

全同态加密(FHE)允许服务器直接在加密的用户提示上运行语言模型,但当前方法速度仍慢到难以实用。密文仅天然支持加法、乘法和旋转操作,且乘法只能组合到有限深度,之后就需要进行代价高昂的引导操作以继续计算。因此,每个非线性激活都必须用迭代方法近似,每次迭代都会用到乘法。更高的迭代次数能提升精度,但会更快耗尽可用深度,触发更多引导操作,而引导操作是延迟的主要来源。现有方法会在整个模型中统一固定迭代次数,而非针对每个位置的误差容限进行定制。我们提出了同态加密感知训练(HEAT),这是一种微调方法,可使每个非线性激活的迭代次数变为可学习参数,让迭代次数与模型权重在训练过程中协同适配。HEAT 针对任务目标优化迭代次数,使模型能适应推理过程中遇到的近似误差,且无需修改架构或从头重新训练。在加密的 GPT-2 解码任务上,HEAT 将迭代次数减少了 3.1 倍,引导次数减少了 1.6 倍,端到端延迟降低了 1.4 倍,同时还提升了与校准基线的解码一致性。

英文摘要

Fully homomorphic encryption (FHE) allows a server to run a language model directly on encrypted user prompts, but current approaches remain prohibitively slow. Ciphertexts natively support only addition, multiplication, and rotation, and multiplications may be composed only to a bounded depth before a costly bootstrapping operation is required to continue. Every nonlinearity must therefore be approximated by an iterative method; each iteration increasing the number of multiplications. A higher iteration count buys precision but exhausts the available depth more frequently and thus triggers more bootstraps, which dominate latency. We introduce Homomorphic Encryption-Aware Training (HEAT), a fine-tuning method that makes the per-nonlinearity iteration counts learnable, enabling them and the model weights to co-adapt during training. HEAT optimizes iterations with respect to the task objective, allowing the model to adapt to approximation errors encountered during inference without architectural changes or retraining from scratch. We further relate iteration count to quantization bit width and bound, at fixed weights, the gap between our objective and quantization-aware training. On encrypted GPT-2 decoding, HEAT reduces iterations by $3.1\times$, bootstraps by $1.6\times$, and end-to-end latency by $1.4\times$, while improving decode agreement over the calibrated encrypted baseline.

发表机构

  • Sapienza University of Rome(罗马大学)
  • George Mason University(乔治梅森大学)
  • Paradigma(帕拉迪格玛公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑