arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

嵌套字节级词汇表:部署成本低,共享成本高——一项预注册的负面结果

Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result

Christos Koutsiaris

arXiv 2608.28151首次发表:更新:

发表机构

SAP P&E(SAP P&E)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究预注册实验发现,嵌套字节级词汇表的切片部署成本低但共享性能不如固定容量模型,控制标记对性能影响小,输出限制成本高,多容量训练可提升鲁棒性且益处源于多粒度训练。

AI 中文摘要

字节级BPE分词器是一组有序的合并规则,因此仅应用前缀会得到一个词汇表,其标记标识符是完整词汇表的前几行。这种前缀嵌套使得一个语言模型可以在多个词汇表大小下运行,使用控制标记指示活跃大小,并通过切片其嵌入和输出头在任何训练大小下部署。我们预注册了五项声明,包括边际、随机种子、对比项和停止规则,并在2亿个标记上训练了30个模型,模型主体参数规模为310万和1060万。切片在数值上是精确的:在76次检查中,切片模型逐位复现了受限完整模型的对数几率,且在不改变延迟的情况下移除了66%的已部署权重。然而,共享模型在32k词汇表大小下与固定容量的专用模型相比,每字节比特数低3.64%(对比1%边际),在8k词汇表大小下低2.96%(对比2%边际)。一项2×2消融实验将控制标记与输出限制分离,发现该标记对性能的影响为+0.07%至+0.13%,所有区间均穿过零,而输出限制的成本为+0.47%至+1.19%;这些因素是替代关系而非互补关系。不过,多容量训练提升了鲁棒性:在印刷噪声下,同一检查点在其精细模式下的性能下降减少了12.5至15.4个点,且在该专用模型的词汇表大小下优于每个固定容量专用模型。既无容量标记也无输出限制的控制组同样鲁棒,表明该益处源于多粒度训练而非条件设置。每个容量的惩罚与该容量训练行的占比相关,为未来工作提供了可证伪的预测。

英文摘要

A byte-level BPE tokenizer is an ordered list of merge rules, so applying only a prefix yields a vocabulary whose token identifiers are the first rows of the full vocabulary. This prefix nesting allows one language model to operate at several vocabulary sizes, use a control token to indicate the active size, and be deployed at any trained size by slicing its embedding and output head. We pre-registered five claims, including margins, seeds, contrasts, and a stop rule, and trained 30 models with 3.1M- and 10.6M-parameter bodies on 200M tokens each. Slicing is numerically exact: across 76 checks, a sliced model reproduces the restricted full model's logits bit for bit and removes 66% of deployed weights without changing latency. However, the shared model trails a fixed-cap specialist by 3.64% bits per byte at 32k against a 1% margin, and by 2.96% at 8k against a 2% margin. A 2x2 ablation separating the control token from output restriction finds that the token changes performance by +0.07% to +0.13%, with all intervals crossing zero, while output restriction costs +0.47% to +1.19%; the factors are substitutes rather than complements. Multi-cap training nevertheless improves robustness: under typographical noise, the same checkpoint degrades 12.5--15.4 points less in its fine mode and outperforms each fixed-cap specialist at that specialist's vocabulary size. A control with neither cap token nor output restriction is equally robust, attributing this benefit to multi-granularity training rather than conditioning. The per-cap penalty tracks each cap's share of training rows, yielding a falsifiable prediction for future work.

Comments5 pages, 2 figures, 4 tables. Pre-registered study. Code and reproducibility materials: https://github.com/unseen1980/captok

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑