发表机构
City University of Hong Kong; Shandong University; Michigan State University(香港城市大学; 山东大学; 密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何在全同态加密下高效部署变压器模型,提出ATLAS自动化框架,通过多目标优化配置每层近似设置,解决近似技术难题,用两阶段优化策略和代理模型应对挑战。
AI 中文摘要
全同态加密(FHE)为隐私推理提供了强大的加密保证,但在FHE下部署变压器模型仍然成本过高。一个关键瓶颈是,诸如softmax、归一化和激活等非线性操作必须用与CKKS方案兼容的多项式近似来代替,而这些近似所消耗的乘法深度主导了推理成本。最近的框架改进了近似技术,但都依赖于手动配置的近似超参数,并且在所有层上统一应用。虽然方便,但这种统一配置方法过于僵化。我们提出了ATLAS,一个通过将问题表述为对延迟和预测准确性的多目标优化来配置每层近似设置的自动化框架。由此产生的问题本质上很困难,ATLAS通过两阶段优化策略和代理模型来加速评估,从而解决这些挑战。
英文摘要
Fully homomorphic encryption (FHE) lets a server run inference on encrypted data with strong privacy guarantees, but running a Transformer under FHE is expensive. Its non-linear operations, such as softmax, normalization, and activation, must be replaced with polynomial approximations that the CKKS scheme supports, and the depth of these approximations dominates inference cost. Existing FHE Transformers use hand-tuned approximation settings, such as iteration count and polynomial degree, applied uniformly across layers, models, and tasks. Hand-tuning is slow and error-prone. Even a single uniform setting has about $10^7$ choices, and manual search cannot exploit layer-wise variation. AutoFHE, the only automated method with multi-objective search, targets ReLU-only CNNs and needs full fine-tuning per candidate, which is too costly for Transformers. Per-layer settings also push the search space to about $10^{85}$ for BERT and ViT and $10^{228}$ for LLaMA3, beyond both manual and fine-tuning-based search. We present ATLAS, a training-free framework that automates this search by treating each layer's approximation setting as a multi-objective optimization over latency and accuracy. The problem is hard: the decision space is large (96 or 256 variables), each configuration takes 70 to 1,000 seconds to evaluate even in cleartext, and 85 to 90 percent of configurations are invalid. ATLAS handles this with a two-stage optimization strategy and a surrogate model, completing the search in about one hour. Compared to an iterative softmax baseline, ATLAS cuts multiplicative depth and end-to-end latency by about 35 percent with little accuracy loss, and works across encoder-only, decoder-only, and vision Transformers, complementing parallel work on packing and matrix multiplication.
CommentsCode: https://github.com/jianhayes/ATLAS