arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LaMoC:面向大语言模型的损失感知模块化压缩方法

LaMoC: Loss-Aware Modular Compression for LLMs

Mohanad Odema, Jacob Song

arXiv 2608.30226首次发表:更新:

发表机构

LG Electronics USA(美国LG电子公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出损失感知模块化压缩方法LaMoC,通过融合激活与经验费舍尔统计优化LLM压缩,在4-8B模型上较最优方法实现困惑度降2.5%、任务精度升1%。

AI 中文摘要

模块化压缩已实现大语言模型(LLMs)的显著参数缩减,同时保留了强大的语言理解能力与下游任务精度。然而,现有联合模块化压缩方法主要依赖激活统计信息,未充分挖掘损失敏感性信息及其模块级表征。本文针对该缺口展开研究,提出LaMoC——一种损失感知模块化压缩方法,通过梯度误差对齐融合激活统计与经验费舍尔(Empirical Fisher)统计。LaMoC通过选择能更好地使局部模块重构误差与下游损失对齐的压缩统计量,提升联合压缩效果。本文贡献有三:(1)将经验费舍尔表征为可与压缩所需激活统计融合的模块级损失感知代理;(2)将联合模块化压缩重新表述为双层优化问题,在调整激活与梯度信息融合率的同时最小化模块重构误差;(3)实现一种经统计验证的实证驱动方法以求解所得压缩问题。本文在涵盖8个模型的4个模型系列上对LaMoC进行评估,在40亿至80亿参数规模的模型上,LaMoC较现有最优模块化压缩方法实现了困惑度平均降低2.5%、任务精度相对提升1%的效果。

英文摘要

Modular compression has enabled considerable parameter reduction in LLMs while preserving strong language understanding and downstream task accuracy. However, existing joint modular compression methods primarily rely on activation statistics, leaving loss-sensitivity information and its module-level characterization underexplored. We investigate addressing this gap with LaMoC, a loss-aware modular compression methodology that blends activation and Empirical Fisher statistics through gradient-error alignment. LaMoC improves joint compression by selecting compression statistics that better align local module reconstruction error with the downstream loss. Our contributions are three-fold: (1) We characterize the Empirical Fisher as a module-level loss-aware proxy that can be blended with the activation statistics required for compression. (2) We reformulate joint modular compression as a two-tiered optimization problem that minimizes module reconstruction error while tuning the activation and gradient information blending rate. (3) We implement an empirically driven methodology with statistical validation to solve the resulting compression problem. We evaluate LaMoC across four model families spanning eight models. On the 4-8B models, LaMoC achieves an average 2.5% reduction in perplexity and a 1% relative improvement in task accuracy over state-of-the-art modular compression methods.

CommentsEMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑