arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

激活感知权重张量化:张量网络LLM压缩的校准期预处理器

Activation-Aware Weight Tensorization: A Calibration-Time Preconditioner for Tensor-Network LLM Compression

Alessandro Beatini, Marco Maronese, Emanuele Rodolà

arXiv 2610.10085首次发表:更新:

发表机构

Sapienza University of Rome(罗马第一大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出激活感知权重张量化(AWT),通过在张量网络压缩前用激活尺度预处理权重矩阵,在多种LLM上以2-6倍压缩率显著提升函数保真度,且无需修改分解求解器。

AI 中文摘要

训练后张量网络压缩用张量列(TT)或树张量网络(TTN)算子替换Transformer线性层,但标准分解最小化的是权重空间中的Frobenius误差,而非层激活分布下的函数误差。我们提出激活感知权重张量化(AWT),这是一种免训练的校准包装器,在不变的TT/TTN求解器之前,用对角激活导出尺度对每个权重矩阵进行预处理,并仅通过输入侧逐元素重缩放来部署结果。在Llama 3.1 8B、Ministral 8B和Qwen2.5 7B上,AWT在2-6倍压缩下持续改进vanilla TT/TTN张量化:在单算子替换下,AWT在三个模型家族和2-6倍压缩设置中,将WikiText困惑度与密集基线的差距缩小了12-35%;而在多算子Llama后缀替换下,在注意力组和全七矩阵设置中,它缩小了27-60%的差距。这些增益也迁移到下游的HellaSwag和ARC-Challenge评估。我们进一步表明,对角预处理是鲁棒性与模块化之间的权衡,而非对角协方差假设:密集的全协方差oracle在80/81种情况下赢得其自身的加权目标,但对角AWT在53/81种情况下提供更好的留出函数保真度。这些结果共同将AWT定位为一种原则性、模块化的预处理器,用于在不修改分解求解器的情况下提高固定TT/TTN压缩流水线的函数保真度。

英文摘要

Post-training tensor-network compression replaces Transformer linear layers with Tensor Train (TT) or Tree Tensor Network (TTN) operators, but standard decompositions minimize weight-space Frobenius error rather than functional error under the layer's activation distribution. We propose Activation-aware Weight Tensorization (AWT), a training-free calibration wrapper that preconditions each weight matrix with a diagonal activation-derived scale before an unchanged TT/TTN solver and deploys the result with only an input-side elementwise rescaling. Across Llama 3.1 8B, Ministral 8B, and Qwen2.5 7B, AWT consistently improves vanilla TT/TTN tensorization at 2-6 times compression: under single-operator replacement, AWT closes 12-35% of the WikiText perplexity gap to the dense baseline across the three model families and 2-6 times compression settings; while under multi-operator Llama suffix replacement it closes 27-60% across attention-group and all-seven-matrix settings. The gains also transfer to downstream HellaSwag and ARC-Challenge evaluations. We further show that diagonal preconditioning is a robustness-modularity tradeoff rather than a diagonal-covariance assumption: a dense full-covariance oracle wins its own weighted objective in 80/81 cases, yet diagonal AWT gives better held-out functional fidelity in 53/81 cases. Together, these results position AWT as a principled, modular preconditioner for improving functional fidelity in fixed TT/TTN compression pipelines without modifying the decomposition solver.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑