发表机构
University of Manitoba; McGill University; Simpleway; The University of Hong Kong; McMaster University(曼尼托巴大学; 麦吉尔大学; Simpleway; 香港大学; 麦克马斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出 FTC 序贯结构化压缩框架,利用 Q/K/V 头结构并适配当前模型,无需微调,在多个 LLM 上以更低困惑度实现高效注意力压缩。
AI 中文摘要
大语言模型(LLM)注意力的训练后压缩常被建模为独立的矩阵近似,忽略了注意力投影间的共享结构以及先前压缩引入的表示偏移。我们提出 FTC,一种序贯结构化压缩框架,在固定存储预算下,将近似自适应于当前压缩模型,同时联合利用原生 Q/K/V 头结构。输出投影被单独处理,以考虑注意力后表示的变化。FTC 无需微调或基于梯度的恢复。在七个参数量从 6B 到 32B 的仅解码器 LLM 上,FTC 在五个现代 GQA 模型的每个测试保留比率下,均取得了所比较方法中最低的 WikiText-2 困惑度,且在激进压缩下收益最大。这些改进可迁移至下游任务,并在 32B 规模下依然显著。
英文摘要
Post-training compression of LLM attention is often formulated as independent matrix approximation, ignoring both the shared structure among attention projections and the representation shift introduced by earlier compression. We propose FTC, a sequential structured compression framework that adapts the approximation to the current compressed model while jointly exploiting the native Q/K/V head structure under a fixed storage budget. The output projection is handled separately to account for the changed post-attention representation. FTC requires neither fine-tuning nor gradient-based recovery. Across seven decoder-only LLMs from 6B to 32B parameters, FTC achieves the lowest WikiText-2 perplexity among the compared methods at every tested keep ratio on five modern GQA models, with the largest gains under aggressive compression. The improvements transfer to downstream tasks and remain substantial at the 32B scale.