arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29397cs.CL

基线形状决定结论:60K参数下三元语言模型的受控再检验

Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters

Gautam Veldanda

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在60K参数下受控重验三元语言模型,发现基线形状主导性能差异,路由模型优势部分源于基线效应,且三元惩罚与训练计划受架构和调优影响。

中文摘要 AI 辅助

三元(1.58比特)权重对微控制器级语言模型具有吸引力,但低于100万参数的范畴主要依赖孤立的单种子比较。一个突出的例子报告称,一个路由三元块(由逐令牌路由器混合的卷积、对角SSM和稀疏注意力)在60K参数下比参数匹配的全精度Transformer高出22%,并将此归因于归纳偏置。我们在一个固定配方下重新运行,每个单元三个种子,在一台笔记本电脑上进行98次字节级运行。(i)基线形状占主导:在1600万字节预算下,参数匹配的Transformer仅通过深度/宽度选择就在验证损失上相差22.6%——远超我们在该处测量的任何架构效应——且形状最佳的Transformer与路由模型持平,因此已发表的差距至少部分归因于基线形状效应;形状的排序随预算而逆转,因此没有单一固定形状可被信任。(ii)在1.3亿字节下,路由模型确实获胜,比我们评估的三种Transformer形状高出22.2%至24.0%——但一个普通的门控对角SSM块进一步比它高出9.1%,且路由模型自身的路由器将其大部分权重放在循环路径上,因此增益并不需要路由。(iii)在较大预算下,三元惩罚因架构而异(最佳Transformer为+5.3%,路由为+19.5%,门控SSM为+28.1%),但我们不能将其仅归因于架构:我们的Transformer以全精度保留学习到的位置嵌入,占其参数的11%至22%,因此它们比所比较的模型量化程度更低。(iv)90/10全精度后三元训练计划优于全三元训练,但仅在阶段2学习率约为预训练峰值10倍时;在常规微调速率下,它看起来差15.3%,从而逆转了结论。从头开始的基线本身未进行学习率调优,这限制了(iii)和(iv)的解释。代码和运行日志已发布。

英文摘要

Ternary (1.58-bit) weights are attractive for microcontroller-class language models, but the sub-1M-parameter regime rests mainly on isolated, single-seed comparisons. One prominent example reports that a routed ternary block (convolution, diagonal SSM and sparse attention mixed by a per-token router) beats a parameter-matched full-precision transformer by 22% at 60K parameters, attributing this to inductive bias. We re-run it under one fixed recipe, three seeds per cell, 98 byte-level runs on one laptop. (i) Baseline shape dominates: at a 16M-byte budget, param-matched transformers span 22.6% in validation loss purely by depth/width choice - far more than any architecture effect we measure there - and the best-shaped transformer ties the routed model, so the published margin is at least partly a baseline-shape effect; the ordering of shapes reverses with budget, so no single fixed shape can be trusted. (ii) At 130M bytes the routed model does win, by 22.2-24.0% over the three transformer shapes we evaluate there - but a plain gated diagonal-SSM block beats it by a further 9.1%, and the routed model's own router puts most of its weight on its recurrent pathway, so the gain does not require routing. (iii) The ternary penalty differs by architecture at the larger budget (+5.3% best transformer vs. +19.5% routed, +28.1% gated SSM), but we cannot attribute that to architecture alone: our transformers keep learned positional embeddings in full precision, 11-22% of their parameters, so they are less quantized than the models they are compared with. (iv) A 90/10 full-precision-then-ternary schedule beats all-ternary training, but only at a stage-2 learning rate about 10x the pretraining peak; at a conventional fine-tuning rate it looks 15.3% worse, reversing the conclusion. The from-scratch baseline was not itself learning-rate tuned, which bounds (iii) and (iv). Code and run logs released.

发表机构

  • Independent Researcher(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑