发表机构
HSE University(高等经济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种可扩展的克罗内克-费舍尔近似方法,实现了十亿参数语言模型的高效海森分析,揭示了值投影层的高敏感性与强跨层相关性,可用于识别大模型脆弱组件,支撑多种压缩优化策略。
AI 中文摘要
本文提出一种可扩展的基于克罗内克的近似方法,该方法无需存储完整费舍尔矩阵即可捕获跨层交互,为无法进行全量计算的十亿参数网络提供了可行的海森分析方案。该方法揭示了一致的脆弱性模式:在多种模型家族中,值投影层表现出最高的敏感性和最强的跨层相关性,而其他组件则表现出特定于架构的行为。通过在量化、稀疏化、层间损坏以及损坏后微调上开展大量实验,我们证明该近似与性能下降和恢复均具有强相关性。我们的框架提供了一种实用、有理论依据的工具,用于识别大模型中的脆弱组件,为指导压缩与优化策略开辟了新途径,例如混合精度分配、逐层稀疏性以及跨层乃至单个权重组的自适应低秩分解。
英文摘要
In this paper, we propose a scalable Kronecker-based approximation that captures cross-layer interactions without storing the entire Fisher matrix, enabling practical Hessian analysis for billion-parameter networks where full computation is infeasible. Our approach reveals consistent vulnerability patterns: value projection layers exhibit the highest sensitivity and strongest cross-layer correlations across multiple model families, while other components exhibit architecture-specific behaviors. Through extensive experiments on quantization, sparsification, inter-layer corruption, and post-corruption fine-tuning, we demonstrate that our approximation strongly correlates with both performance degradation and recovery. Our framework provides a practical, theoretically grounded tool for identifying fragile components in large models, opening new avenues for guided compression and optimization strategies, such as mixed-precision allocation, layer-wise sparsity, and adaptive low-rank decomposition across layers and even individual weight groups.