arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.01776cs.CLcs.AIcs.CV

FreeAct: 为LLM量化释放激活

FreeAct: Freeing Activations for LLM Quantization

  • National University of Singapore(新加坡国立大学)
  • Huawei Technology(华为技术有限公司)
  • Central South University(中央南大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiaohao Liu, Xiaobo Xia, Manyi Zhang, Ji-Fu Li, Xianzhi Yu, Fei Shen, Xiu Su, See-Kiong Ng, Tat-Seng Chua

更新

AI总结:

FreeAct通过放松静态转换约束,为LLM量化提供动态适应的激活转换方法,实现性能提升5.3%。

AI中文摘要:

量化对于缓解大型语言模型(LLM)显著的内存和计算开销至关重要。尽管新兴的基于转换的方法通过使用正交矩阵将特征空间投影到更平滑的流形上,成功地增强了量化效果,但它们通常强制实施一种刚性的一对一转换约束。这种静态方法无法考虑输入激活中的动态模式,特别是在扩散LLM(dLLM)和多模态LLM(MLLM)中,不同类型的token表现出不同的分布。为此,我们提出了FreeAct,一种新的量化框架,它放松了静态的一对一约束,以适应动态的激活差异。理论上,我们利用激活的秩不足特性,推导出一个扩展超过简单逆矩阵的解决方案空间,使激活转换与权重解耦。方法上,FreeAct识别token特定的动态(即视觉vs文本,或掩码token)并为激活侧分配不同的转换矩阵,同时保持权重的统一静态转换。在dLLM和MLLM上的广泛实验表明,FreeAct在基线上显著优于,性能提升高达5.3%,并进行了深入分析。我们的代码将公开发布。

英文摘要:

Quantization is pivotal for mitigating the significant memory and computational overhead of Large Language Models (LLMs). While emerging transformation-based methods have successfully enhanced quantization by projecting feature spaces onto smoother manifolds using orthogonal matrices, they typically enforce a rigid one-to-one transformation constraint. This static approach fails to account for the dynamic patterns inherent in input activations, particularly within diffusion LLMs (dLLMs) and Multimodal LLMs (MLLMs), where varying token types exhibit distinct distributions. To advance this, we propose FreeAct, a novel quantization framework that relaxes the static one-to-one constraint to accommodate dynamic activation disparities. Theoretically, we leverage the rank-deficient nature of activations to derive a solution space that extends beyond simple inverse matrices, enabling the decoupling of activation transformations from weights. Methodologically, FreeAct identifies token-specific dynamics (i.e., vision v.s. text, or masked tokens) and allocates distinct transformation matrices to the activation side, while maintaining a unified, static transformation for the weights. Extensive experiments across dLLMs and MLLMs demonstrate that FreeAct significantly outperforms baselines, up to 5.3% performance improvement, with in-depth analyses. Our code will be publicly released.

补充信息

↑