LaPrune:百万规模下的可控可微稀疏性
LaPrune: Controllable Differentiable Sparsity at Million Scale
浏览论文内容
中文总结 AI 辅助
本研究提出LaPrune算法,通过LapSum屏障与归一化二阶矩约束实现百万规模下可控可微稀疏性,推导了相关理论保证,解决了稀疏模型训练中梯度与质量耦合的问题。
中文摘要 AI 辅助
Top-$k$选择决定稀疏模型的哪些组件保持激活状态:硬选择会阻断梯度,而连续松弛通常将掩码的硬度与所选质量耦合。我们提出LaPrune,这是一种数学上精确预算的可微层,在保留所选质量的同时控制归一化二阶矩。LapSum屏障用于保留选择质量,归一化二阶矩约束使掩码从密集等质量分配向每次预算下的硬Top-$k$移动。我们推导了饱和分数的总体预测、近二元极限定律,以及近零分数的严格最坏情况保证。归一化硬度参数对分数尺度不变,而固定LapSum温度则不具备该特性。
英文摘要
Top-$k$ selection determines which components of a sparse model remain active. Hard selection blocks gradients, while continuous relaxations often couple mask hardness to the selected mass. We introduce LaPrune, a mathematically exact-budget differentiable layer that controls the normalized second moment while preserving the selected mass. A LapSum barrier preserves the selection mass, and a normalized second-moment constraint moves the mask from a dense equal-mass allocation toward hard top-$k$ at each budget. We derive a population prediction of the saturated fraction, a near-binary limiting law, and a tight worst-case guarantee on the near-zero fraction. The normalized hardness parameter is invariant to score scale, while a fixed LapSum temperature is not.
发表机构
- Wrocław University of Science and Technology(弗罗茨瓦夫理工大学)
- Faculty of Mathematics and Computer Science, Jagiellonian University(雅盖隆大学数学与计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。