arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过规范校正通道移除实现Transformer的密集结构压缩

Dense Structural Compression of Transformers via Gauge-Correct Channel Removal

Jed A. Duersch, Naïm Es-Sebbani, Nathanaël Haas, Zied Bouraoui

arXiv 2609.07264首次发表:更新:

发表机构

CNRS; CRIL(法国国家科学研究中心; CRIL研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出GaugeLasso方法,通过规范校正通道移除在训练中动态压缩Transformer结构,实现148-255倍计算压缩并保持精度,揭示持续压力对发现高效模型的关键作用。

AI 中文摘要

每个token的推理能耗决定了已部署Transformer的成本和碳足迹。该能耗主要由密集矩阵乘积主导,这些乘积涉及融合乘加(FMA)运算和内存流量。为了在保持密集张量以获得高GPU吞吐量的同时减少这些计算,我们从第一性原理出发开发了一种方法论,在训练过程中调整结构复杂度,以最大化每单位计算的推理效用。通道惩罚将整个张量切片驱动至零,从而在保持密度和网络功能的同时实现物理移除。自然的方法是惩罚通过每个通道作用的算子组件的范数,但这种方法被证明会因规范自由度而不稳定。我们通过GaugeLasso解决了这一病理问题:当网络收敛到规范平衡时,加性对称组套索惩罚恢复为乘积范数的单调函数。我们的均衡分析实现了逐通道校准,以正确抑制在每单位计算推理效用方面表现不佳的切片。在自适应压力下,网络重组为深度相关的结构轮廓,这些轮廓可能远小于学习任务所需的架构。在$\mathbb{F}_{31}$上的多项式长除法中,计算压缩了148到255倍,且精度完美。在字符级语言建模中,压缩模型在相同FMA下优于手工设计的基线。在掩码自编码中,一次压缩试验揭示了哪些轴过度配置、哪些轴饱和,从而指导了更好的第二次设计。压缩还随着模型的进展单调地加速训练。使用相同效用排名的事后剪枝无法达到这些结构,表明持续压力是发现高效模型的核心。重新训练一个发现的架构在我们的统计任务上恢复了基线质量,但在精确算法任务上失败。

英文摘要

Inference energy per token drives the cost and carbon footprint of deployed transformers. It is dominated by dense matrix products that incur fused multiply-accumulate (FMA) operations and memory traffic. To reduce these computations while retaining dense tensors for high GPU throughput, we develop a methodology from first principles to adapt structural complexity during training to maximize inference utility per unit compute. Channel penalties drive entire tensor slices to zero to enable physical removal while preserving density and the network function. The natural approach, penalizing the norm of operator components acting through each channel, is provably destabilized by gauge freedom. We resolve this pathology with GaugeLasso: additive symmetric group-lasso penalties that recover a monotone function of product-norms when the network converges to gauge balance. Our equilibrium analysis enables per-channel calibration to correctly suppress slices that under-perform in inference utility per unit compute. Under adaptive pressure, the network reorganizes into depth-dependent structural profiles that can be far smaller than the architecture required to learn the task. On polynomial long division over $\mathbb{F}_{31}$, compute compresses from 148 to 255 times with perfect accuracy. On character-level language modeling, compressed models outperform the hand-designed baseline at equal FMA. On masked autoencoding, a compression trial exposes which axes were over-provisioned and which saturated, guiding a better second design. Compaction also accelerates training monotonically as the model progresses. Post-hoc pruning with the same utility ranking cannot reach these structures, showing that sustained pressure is central to discovery of efficient models. Retraining a discovered architecture recovers baseline quality on our statistical tasks, but fails on our exact algorithmic task.

Comments10 pages main text, 28 pages total including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑