arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06974cs.CLcs.AI

训练时过完备,部署时紧凑:扩展结构化大语言模型剪枝的恢复能力

Train Overcomplete, Deploy Compact: Scaling Recovery Capacity for Structured LLM Pruning

Seungmin Oh, Donggeon Lee, Jongbin Ryu

首次发表
浏览论文内容

中文总结 AI 辅助

针对结构化剪枝恢复阶段的能力-知识不对称问题,提出过完备重参数化框架OverRep,通过训练时过参数化并代数合并为紧凑模块,在保持推理成本的同时显著提升剪枝后模型的推理性能。

中文摘要 AI 辅助

大语言模型在各类任务上表现出强大的性能,但由于内存、延迟和能耗需求,其部署成本仍然高昂。结构化剪枝通过移除架构组件来降低这些成本,但其恢复阶段常常受到恢复模块的表征能力与被移除知识复杂度之间不匹配的限制。我们将这一瓶颈称为“能力-知识不对称”,并提出了OverRep,一种用于结构化大语言模型剪枝的过完备重参数化框架。遵循“训练时过完备,部署时紧凑”的原则,OverRep在训练期间临时对恢复模块进行过参数化,以吸收从原始模型中蒸馏出的复杂知识。恢复后,过完备的重参数化通过代数方式合并为数学上等价的紧凑模块,从而保持剪枝模型推理时的架构和计算成本。OverRep还引入了一种退火激活,使得在收敛到线性机制以进行精确代数合并的同时,能够实现非线性训练动态。在三个骨干家族中,与强恢复基线相比,OverRep在25%和50%剪枝率下分别将保留的推理性能提升了最多5.5和8.4个百分点,同时内存使用和TFLOPs与现有恢复方法相当。我们的代码可在以下https URL获取。

英文摘要

Large language models achieve strong performance across diverse tasks, but deployment remains costly because of memory, latency, and energy demands. Structured pruning reduces these costs by removing architectural components, yet its recovery stage is often limited by a mismatch between the recovery module's representational capacity and the complexity of the removed knowledge. We call this bottleneck the capacity-knowledge asymmetry and propose OverRep, an Overcomplete Reparameterization framework for structured LLM pruning. Following the principle of "train overcomplete, deploy compact", OverRep temporarily overparameterizes the recovery module during training to absorb complex knowledge distilled from the original model. After recovery, the overcomplete re-parameterization is algebraically merged into a mathematically equivalent compact module, preserving the pruned model's inference-time architecture and computational cost. OverRep further introduces an annealed activation that enables nonlinear training dynamics while converging to a linear regime for exact algebraic merging. Across three backbone families, OverRep improves retained reasoning performance over strong recovery baselines by up to 5.5 and 8.4 points at 25% and 50% pruning, respectively, while keeping memory usage and TFLOPs comparable to existing recovery methods. Our code is available at https://github.com/mmai-laboratory/OverRep.

发表机构

  • Ajou University(明知大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑