AI 中文总结
该研究以 Pythia 模型家族为例,揭示 Transformer 尺寸间转换的关键在于初始化,提出最小二乘补偿与方差保持缩放两个控制项,在低预算下可高效实现模型转换,优于从头训练与子克隆方法。
AI 中文摘要
模型家族会从头训练每一种尺寸的模型。预训练的大模型能否被转换为其小型同类模型?我们对 Pythia 模型家族中 14 亿参数(1.4B)模型到 4.1 亿参数(410M)模型的转换进行了端到端的表征:(i)不同尺寸模型的表征对齐程度很高(岭回归 R²=0.84),而参数对齐程度较弱;(ii)稠密权重投影在功能上具有破坏性——这并非组装的人为结果,因为基混合会破坏旋转位置编码、多头结构、GELU 激活函数和 LayerNorm 结构;(iii)在最佳拟合线性算子之后,权重残差在洗牌控制下与噪声在统计上无法区分;(iv)因此,转换的价值存在于初始化中。在匹配预算的持续预训练中,我们将转换分解为两个独立的控制项——最小二乘补偿(功能:最佳零样本性能)和方差保持缩放(动态:最佳端点性能)。补偿是一种 token 高效、低预算的改进,而非通用改进:在 3000 万 token 时,它在宽度缩减对(84.0 ± 1.8 对比 89.7 ± 3.7,3 次实验均成功)和保留的深度缩减对(109.3 对比 117.9,3 次实验均成功)上均优于最强的子克隆变体,能用更少的 token 达到给定质量;在 33 倍更大的预算下,两者收敛至同等水平(40.0 对比 40.0),均远优于从头训练的模型,而迁移初始化始终优于从头训练模型——在低预算时优势可达 18 倍,在收敛和最大规模时差距缩小。我们进一步划定了该方法的边界:在约 5 倍于源模型规模(69 亿参数(6.9B)模型到 14 亿参数(1.4B)模型)时,同时使用两个控制项会出现过度校正,我们将其归因于大宽度下补偿求解的病态条件,这表明维度感知正则化是解决方法。代码、检查点和冻结评估语料已公开。
英文摘要
Model families are trained size by size. Can a pretrained large model instead be converted into a smaller sibling? We study the 1.4B->410M conversion in Pythia end to end. Representations align strongly across sizes (ridge R^2=0.84); parameters align weakly. Dense weight projection is destructive; a bit-exact control places the fault in basis mixing, which breaks rotary, per-head, GELU, and LayerNorm structure. Residuals after the best-fit linear operator carry no learnable or transferable signal under shuffle controls, so conversion value lives in initialization. Matched-budget continued pre-training separates two independent levers: least-squares compensation (function lever, best zero-shot) and variance-preserving rescale (dynamics lever, best endpoints). Placement follows the architecture: compensation is well-posed exactly where no normalization sits between cut and read; norm-fronted paths take rescale. Compensation is a low-budget, token-efficiency win, not a universal one. At 30M tokens it beats the best subcloning variant on a width-reduced pair (84.0+-1.8 vs. 89.7+-3.7, 3/3 seeds) and a held-out depth-reduced pair (109.3 vs. 117.9, 3/3 seeds). Selection given the same activation statistics recovers under half of that gap (3/3 seeds): the gain is the re-fit, not the information. At 33x the budget the two reach parity (40.3+-0.3 vs. 40.3+-0.5, 3 seeds), both far ahead of from-scratch, which transfer always beats (up to 18x at low budget, narrowing at convergence and at the largest scale). At ~5x the donor scale (6.9B->1.4B) stacking both levers over-corrects, consistent with an ill-conditioned compensation solve at large width, pointing to dimension-aware regularization as a fix. The init also beats structured pruning plus distillation, the standard pipeline, at matched budget, and improves further combined with it. Code, checkpoints, and the frozen eval corpus are released.
Commentsv3: 3-seed 1B convergence and extended evaluation, information-matched selection control, LayerNorm/structure decomposition of the projection failure, 3-seed distillation comparison; retitled. 18 pages, 4 figures, 13 tables