Fiona:利用打包感知三元权重加速FHE推理
Fiona: Accelerating FHE Inference with Packing-Aware Ternary Weights
浏览论文内容
中文总结 AI 辅助
FIONA是一种离线优化器,通过选择性三元化权重并保留敏感组全精度,在打包布局下减少PMult操作,加速FHE推理,在三个模型上实现1.68-2.38倍加速且精度损失小于1%。
中文摘要 AI 辅助
全同态加密(FHE)能够直接在加密输入上进行神经网络推理,但其速度仍比明文推理慢数个数量级。将服务器的明文权重应用于加密激活涉及明文-密文乘法(PMult),在近期系统中占推理时间的一半以上。三元量化可以将这些乘法替换为加法和减法,但在打包执行下,这种节省很少实现。单个PMult应用于由打包布局固定的权重组,仅当该组中所有权重共享相同三元值时才可避免。然而,对所有组进行三元化会大幅降低精度。我们提出FIONA,一种离线优化器,根据三元转换对模型性能的估计影响,在给定打包布局内选择性三元化权重。FIONA鼓励每个权重组内共享三元值,并保留敏感组的全精度权重,因此三元化和全精度路径在同一层内共存。然后,它精确编译这些混合算子,将公共缩放因子一次性应用于累积输入,并在输出间重用和。权重三元化还可以缩小下游多项式的输入范围。FIONA在累积精度预算下拟合低阶替代,减少乘法深度和自举。在VGG11、ViT和BERT上,FIONA分别减少PMult操作53.4%-79.5%,并将端到端加密推理加速2.38倍、1.68倍和1.84倍,三个模型的精度损失均小于1%。
英文摘要
Fully homomorphic encryption (FHE) enables neural network inference directly on encrypted inputs, but it remains orders of magnitude slower than plaintext in- ference. Applying the server's plaintext weights to encrypted activations involves plaintext-ciphertext multiplications (PMult) and accounts for more than half of inference time in recent systems. Ternary quantization can replace these multipli- cations with additions and subtractions, but the savings rarely materialize under packed execution. A single PMult applies a weight group fixed by the packing layout and can be avoided only when all its weights share the same ternary value. Ternarizing all groups, however, largely degrades accuracy. We present FIONA, an offline optimizer that selectively ternarizes weights within a given packing layout based on the estimated effect of ternary conversion on the model's performance. FIONA encourages a shared ternary value within each weight group and retains full-precision weights for sensitive groups, so ternar- ized and full-precision paths coexist within a layer. It then compiles these hybrid operators exactly, applying common scaling factors once to accumulated inputs and reusing sums across outputs. Weight ternarization can also narrow the input ranges of downstream polynomials. FIONA fits lower-degree replacements under a cumulative accuracy budget, reducing multiplicative depth and bootstrapping. On VGG11, ViT, and BERT, FIONA reduces PMult operations by 53.4-79.5% and accelerates end-to-end encrypted inference by 2.38x, 1.68x, and 1.84x, re- spectively, with less than 1% accuracy loss across all three models.
发表机构
- The Hong Kong University of Science and Technology(香港科技大学)
- State Key Laboratory of Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。