发表机构
Incept Labs; Titan Holdings(因塞普特实验室; 泰坦控股公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对低秩克隆(LRC)蒸馏中MLP存在的可达性差距,提出训练部署权重的原则,通过两种可合并实现恢复闲置容量,在不同教师模型上取得显著性能提升,大幅降低蒸馏所需token数。
AI 中文摘要
压缩后的学生模型存在两种无需一致的权重形态:推理时部署的权重,以及其训练可触及的权重族。本文表明,最先进的权重继承蒸馏器低秩克隆(Low-Rank Clone, LRC)会部署全宽度的学生多层感知机(MLP),但将训练绑定到教师诱导的切片上,导致每个部署矩阵的62.5%-81.4%独立线性自由度无法触及——这些自由度在推理时占用资源,却从未参与训练。本文提出的原则仅一句话:训练你部署的模型。从相同的LRC预热启动开始,本文将训练对象设为整个部署矩阵,无需改变部署形态、部署参数数量或推理浮点运算次数(FLOPs),通过两种可合并实现(Dense-LRC和CORE-LRC),两者最终都合并为一个部署权重。这恢复了闲置容量:针对三位教师模型(Llama3.2-3B、Llama3.1-8B、Qwen2.5-3B),取每位教师对应的更强实现,在匹配预算的普通LRC基线之上,平均9项指标提升+2.36/+2.71/+10.45;在最宽的教师模型Qwen上,提升最大,其在10B token时达到原方案约20B token的准确率(token效率提升1倍),其中严格同谱系分支仍提升+6.39,这是完全可控的数值。对照实验强烈支持将提升归因于可触及集合的扩大,而非额外参数或训练方案。从约10B蒸馏token加简短的监督微调(SFT),参数减半的1.5B学生模型在评估噪声范围内匹配其约9T token教师的9任务宏平均,仅残留MMLU deficit;2.7B学生模型击败Meta官方对Llama3.1-8B的压缩版本,而压缩token数少约900倍(该token数低于未匹配方案,并非计算量声明)。所有结果均为LRC骨干上的单种子运行结果。
英文摘要
A compressed student has two shapes that need not agree: the weight it deploys at inference and the weight family its training can reach. We show that a state-of-the-art weight-inheritance distiller, Low-Rank Clone (LRC), deploys a full-width student MLP but ties training to a teacher-induced slice, leaving 62.5-81.4% of each deployed matrix's independent linear degrees of freedom unreachable-paid for at inference, never trainable. Our principle is one line: train what you deploy. From the identical LRC warm start, we make the training object the entire deployed matrix, with no change in deployed shape, deployed parameter count, or inference FLOPs, via two mergeable realizations (Dense-LRC and CORE-LRC) that both collapse to one deployed weight. This recovers stranded capacity: taking the stronger realization per teacher, +2.36/+2.71/+10.45 Avg9 over matched-budget plain-LRC baselines across three teachers (Llama3.2-3B, Llama3.1-8B, Qwen2.5-3B), with the largest gain on the widest teacher (Qwen), where it reaches the original recipe's approx. 20B-token accuracy at 10B tokens (2x token efficiency); there the strictly same-lineage arm still recovers +6.39, the fully controlled figure. Controls strongly support attributing the gain to the enlarged reachable set, rather than to added parameters or the recipe. From approx. 10B distillation tokens plus a short SFT, a half-parameter 1.5B student matches its approx. 9T-token teacher's 9-task macro-average, within evaluation noise and with a residual MMLU deficit, and a 2.7B student beats Meta's own official compression of Llama3.1-8B at ~900x fewer compression tokens (a token count under unmatched recipes, not a compute claim). All results are from single-seed runs on the LRC backbone.