arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

层并行推理减少了Transformer中加密的非线性深度

Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers

Ligong Han, Kai Xu, Hao Wang, Ruijiang Gao, Han Gao, Akash Srivastava

arXiv 2607.04819首次发表:更新:

发表机构

MBZUAI IFM; Red Hat AI Innovation; MIT-IBM Watson AI Lab; Core AI, IBM; University of Texas at Dallas(穆罕默德·本·扎耶德人工智能大学智能未来研究院; 红帽人工智能创新实验室; 麻省理工学院-IBM沃森人工智能实验室; IBM核心人工智能部门; 德克萨斯大学达拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究结构化牛顿层并行性(SNLP)能否使Transformer层间组合更适合全同态加密(FHE),通过基于切比雪夫多项式近似的模拟框架测量误差积累,结果表明SNLP可减少推理步骤并降低误差放大。

AI 中文摘要

全同态加密(FHE)实现加密数据计算,实用的加密Transformer推理受非线性块顺序组合瓶颈限制。研究SNLP能否使层间组合更FHE友好,通过模拟框架测量8个模型和4个架构系列顺序与SNLP推理下的误差积累,结果显示SNLP有优势,且softmax近似主导误差预算。

英文摘要

Fully homomorphic encryption (FHE) enables computation on encrypted data, but practical encrypted Transformer inference is bottlenecked by the sequential composition of many nonlinear blocks. We study whether Structured Newton Layer Parallelism (SNLP) can make this inter-layer composition more FHE-friendly: each Transformer block still requires polynomial approximations for operations such as softmax and RMSNorm, but SNLP reduces the layerwise sequential nonlinear depth from L stages to a small number of solver iterations plus linear structured corrections. Using a simulation framework based on Chebyshev polynomial approximations, we measure error accumulation under sequential versus SNLP inference across 8 models and 4 architecture families. On a 0.5B IDN-trained model, SNLP reduces symbolic bootstraps from 53 to 20 (2.65x) with only +1.2% perplexity degradation, while lowering error amplification (1.36x vs. 1.42x). Across all tested models, SNLP has lower amplification than sequential inference. Ablations show that softmax approximation dominates the error budget and CKKS arithmetic noise is negligible in our setting, suggesting that SNLP is complementary to block-level FHE-friendly operator design rather than a replacement for it.

CommentsCode is available at https://github.com/phymhan/nanochat-snlp/tree/snlp-fhe

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑