arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14715cs.AIcs.CL

亚1.5亿参数规模下的深度与尺度:JugnuLM-53M 对比 JugnuLM-110M

Depth and Scale in the Sub-150M Regime: JugnuLM-53M vs JugnuLM-110M

Dushyant Rajput, Nirdesh Chauhan, Siddharth Kosaraju

首次发表
浏览论文内容

中文总结 AI 辅助

将常规预训练方案从53.5M扩展到110M参数,采用深层薄架构,在更少数据下提升性能,并通过消融实验确定Muon优化器为最佳改进。

中文摘要 AI 辅助

我们将常规的亚1.5亿参数预训练方案从53.5M参数扩展到109.7M参数,保持方法不变(采用Qwen3风格的解码器,包含分组查询注意力、RoPE、SwiGLU、RMSNorm、QK-Norm和z-loss;使用FineWeb-Edu数据),仅将几何结构改为深而薄的23层×576隐藏维度设计。更大的模型在各项指标上均有提升——BLiMP从78.1提升至81.3,ARC-Easy从51.4提升至52.5,WikiText-2字节困惑度从2.04降至1.95——其81.3%的BLiMP成绩基本与GPT-X2-125M(81.28)持平,而参数减少了约12%。值得注意的是,110M模型在更少的训练token(约8B对比12B)下实现了这一成绩,因此提升归因于容量和深度,而非更多数据。两个模型均刻意采用常规设计;本报告是一个干净的缩放对照实验,也是该规模下进一步改进模型消融研究的基线层级(R0)。随后是消融阶梯:值残差(R1)和Muon优化器(R2)使ARC-Easy累计提升+3.6(从52.5提升至56.1),同时BLiMP几乎持平,因此被保留;多样数据混合(R3)和两种logit蒸馏设置(R4a/R4b)未被保留——这是诚实的负面结果。R3将ARC-Easy的提升归因于FineWeb-Edu的教育性过滤而非原始多样性;从1.7B教师模型蒸馏可以达到同类领先的ARC-Easy成绩(56.99,与GPT-X2-125M持平),但代价是困惑度增加,而降低KD权重又会消除这一提升——因此R2仍是保留的最佳堆栈。

英文摘要

We scale our conventional sub-150M pretraining recipe from 53.5M to 109.7M parameters, holding the method fixed (Qwen3-style decoder with grouped-query attention, RoPE, SwiGLU, RMSNorm, QK-Norm, and a z-loss; FineWeb-Edu data) and changing only the geometry to a deep-and-thin 23-layer x 576-hidden design. The larger model improves across the board -- BLiMP 78.1 -> 81.3, ARC-Easy 51.4 -> 52.5, WikiText-2 byte-perplexity 2.04 -> 1.95 -- and its 81.3% BLiMP essentially matches GPT-X2-125M (81.28) at about 12% fewer parameters. Notably the 110M model achieves this on fewer training tokens (about 8B vs 12B), so the gain is attributable to capacity and depth, not more data. Both models are deliberately conventional; this report is a clean scaling control and the baseline rung (R0) of an ablation study of what further improves models in this regime. An ablation ladder follows: value residuals (R1) and the Muon optimizer (R2) lift ARC-Easy by a cumulative +3.6 (52.5 -> 56.1) at a near-flat BLiMP and are kept; a diverse data blend (R3) and two logit-distillation settings (R4a/R4b) are not kept -- honest negatives. R3 pins ARC-Easy to FineWeb-Edu's educational filtering rather than raw diversity; distillation from a 1.7B teacher can reach the class-leading ARC-Easy (56.99, matching GPT-X2-125M) but only at a perplexity cost that dialing KD down then erases -- so R2 remains the best kept stack.

发表机构

  • AltSlate Labs LLP

机构由 AI 辅助整理,请以论文原文为准。

↑