arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DegreeSpar:面向高效安全Transformer推理的结构化度数稀疏化

DegreeSpar: Structured Degree Sparsity for Efficient Secure Transformer Inference

Yifei Cai, Zhuoran Li, Xiaozuo Shen, Hongyi Wu, Chunsheng Xin

arXiv 2609.32204首次发表:更新:

发表机构

Iowa State University; University of Arizona(爱荷华州立大学; 亚利桑那大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DegreeSpar通过结构化多项式度数稀疏化统一压缩安全Transformer推理中的非线性计算,实现2.29x-6.63x加速并保持准确率,确立度数作为共享优化空间。

AI 中文摘要

安全Transformer推理保护敏感输入,但会带来巨大的密码学开销,其中Softmax和GeLU等非线性操作成为主要瓶颈。现有压缩方法通过分别定义的压缩变量来降低非线性复杂度、序列相关计算或模型结构。然而,在激进压缩下,这些独立优化的扰动可能累积:在匹配的压缩水平下,叠加代表性的近似、令牌剪枝和模型剪枝方法会将ViT-S的准确率从80.20%降至76.41%。我们提出DegreeSpar,将安全Transformer压缩形式化为非线性多项式度数的结构化稀疏化。多项式度数直接控制安全非线性评估的成本,而计算对齐的零度数结构将令牌级和模型维度计算暴露为可在同一优化空间内移除。DegreeSpar进一步结合了针对低度数Softmax和GeLU的近似感知训练,实现激进的度数降低,并为结构化计算移除创造所需的优化空间。在视觉和语言Transformer上,DegreeSpar在模型规模、任务和序列长度上持续改善准确率-延迟权衡,相较于相应基线实现了2.29倍至6.63倍的加速。在相同网络设置下,DegreeSpar在BERT/SST-2上以110.55秒达到92.68%的准确率,而最接近的先前混合安全推理方法CipherPrune以167.26秒达到92.66%。这些结果确立了结构化多项式度数作为安全Transformer压缩的有效共享优化空间。

英文摘要

Secure Transformer inference protects sensitive inputs but incurs substantial cryptographic overhead, with nonlinear operations such as Softmax and GeLU becoming major bottlenecks. Existing compression methods reduce nonlinear complexity, sequence-dependent computation, or model structure through separately defined compression variables. Under aggressive compression, however, these independently optimized perturbations can accumulate: at a matched compression level, stacking representative approximation, token-pruning, and model-pruning methods reduces ViT-S accuracy from 80.20% to 76.41%. We introduce DegreeSpar, which formulates secure Transformer compression as structured sparsification over nonlinear polynomial degrees. Polynomial degree directly controls the cost of secure nonlinear evaluation, while computation-aligned zero-degree structures expose token-level and model-dimension computation as removable within the same optimization space. DegreeSpar further incorporates approximation-aware training for low-degree Softmax and GeLU, enabling aggressive degree reduction and creating the optimization headroom required for structured computation removal. Across vision and language Transformers, DegreeSpar consistently improves the accuracy-latency trade-off across model scales, tasks, and sequence lengths, achieving speedups from 2.29x to 6.63x over the corresponding baselines. Under the same network setting, DegreeSpar achieves 92.68% accuracy on BERT/SST-2 in 110.55 s, compared with 92.66% in 167.26 s for CipherPrune, the closest prior hybrid secure-inference approach. These results establish structured polynomial degree as an effective shared optimization space for secure Transformer compression.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑