当压缩分数无法做出决策时:群体鲁棒大语言模型剪枝的信息边界
When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning
浏览论文内容
中文总结 AI 辅助
该研究针对压缩统计量无法准确选择群体鲁棒 LLM 剪枝候选的问题,通过信息接口建模偏差,提出粗粒度深度分配等方法,实现最差群体困惑度与 KL 散度降低,提升剪枝效果。
中文摘要 AI 辅助
可复现的压缩统计量仍可能选错候选对象。一项 dense 剪枝分数具有 0.906 的分半信度,曾预测能获得 16.1% 的增益,但该分数选定的端点相较于两个对照组分别差 6.0% 和 7.7%。我们通过信息接口对该偏差进行建模,这些接口界定了每个统计量所能区分的内容。对于等权重群体,二次型定律给出了正线性固定候选损伤的精确池化代价,包括对角和全 PSD 二阶矩。三种双世界构造及精确观测纤维半径,刻画了池化矩、群体局部矩和参考路径曲率未解决的问题。群体解析对角可恢复广泛的损伤排序(斯皮尔曼相关系数 0.9239),但精细排序仍较弱。相较于平衡均匀分配,粗粒度深度分配在三个 dense 大语言模型(LLM)上将最差群体困惑度通胀降低了 12.6%至 20.9%。特定模型的完整掩码端点选择较这些参考方法提升了 2.7%至 8.0%。在 OLMoE 中,路由器迹线可预测单例方向(最强重标记下为 114/192 对 81/192)。对某一层的有限菜单决策可在保留数据上实现最差群体 KL 散度降低 13.7%和 7.2%。局部测量构造候选对象,选择由完整候选端点或经验证的均匀保证许可,且不确定性与每一次比较校准。
英文摘要
A stable compression score can still select the worse model. In our dense study, a split-half reliable path-quadratic score predicted a 16.1\% gain, while the selected endpoints were 6.0--7.7% worse than two controls. We ask what a compression statistic can justify when deployment cares about the worst supplied group. We treat each statistic as an information interface. Its observation leaves a fiber of compatible endpoint-risk tables, and only orders fixed across that fiber are identified. Cone and fiber identities quantify the remaining uncertainty, while matched observations reverse endpoint order for pooled moments, group-local moments, and reference-path curvature. Sequential composition adds one state variable: the slack from each group risk to the current maximum. This vector determines every unrestricted one-step response, and a margin condition keeps the active group fixed along paths with bounded relative drift. The experiments follow the same ladder. Across three dense LLMs, an early-preserving allocation reduces worst-group perplexity inflation by 12.6--20.9%; target-matched complete-menu selection improves over its references by 2.7--8.0%. Across all 16 routed layers of OLMoE, pooled endpoint refresh lowers held-out worst-group teacher KL by 15.8% over the best static score. A compute-matched hard-max trajectory ends 32.7% worse than pooled, and neither adaptive trajectory improves excess NLL. Local evidence can narrow a menu. Complete endpoints rank that menu, while multistep claims also require control of the evolving active face and future candidates.
发表机构
- KTH Royal Institute of Technology(皇家理工学院)
机构由 AI 辅助整理,请以论文原文为准。