AI 中文总结
该研究发现字节级语言模型的熵触发补丁布局遗漏了需计算的关键位置,提出边界依赖信号结合熵的方法,在GSM8K等任务上显著提升了计算结果的准确率。
AI 中文摘要
字节级语言模型(如字节潜变量变换器BLT)将字节分组为补丁,并针对每个补丁运行一次其大型全局模型。BLT在小型模型的下一字节熵较高的位置启动补丁,因此全局计算会分配到下一字节难以预测的地方。我们表明,该规则存在系统性盲点:那些类型可预测但必须计算其值的位置,例如数学解题中“=”后的数字。在严格的补丁预算下,熵触发的布局会跳过这些位置,导致其准确率大幅下降。在Meta的BLT-1B模型中,当10%的字节作为补丁起点时,熵规则在GSM8K解题的计算结果中设置了16%的补丁起点,仅使其中19.0%完全正确;而在相同补丁数量下,将每个“=”后设为边界时准确率达51.8%,将熵与无标签的边界依赖信号结合时准确率达67.1%(默认布局在26%的字节处:76.8%)。该差距在通过低秩微调调整BLT-1B以适应预算时依然存在(32.9% vs 72.7%,每个规则3次运行,配对p<1e-200),且在10%预算下从头训练的字节模型中,差距随模型规模增大而扩大:在100万、1200万和5000万参数时,边界依赖在最终答案上比熵规则分别高出-1.6、+10.1和+19.8个百分点,在5000万参数时,其计算结果准确率达35.9%,而熵规则仅为13.9%(各3个随机种子)。Scratchpad Patching的熵触发在最终答案上与随机scratchpad无显著差异(5.6% vs 5.9%,5个种子),而答案起始scratchpad的准确率达38.1%。该效应仅针对计算值:复制和查找操作收益甚微,模型无法计算的值则无收益。边界依赖(即移除补丁起点时模型自身损失的上升,按两字节上下文测量)无需标签即可识别这些位置:结合熵后,其在计算结果上优于人工编写的规则。
英文摘要
Byte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch. BLT starts a patch where a small model's next-byte entropy is high, so global compute goes where the next byte is hard to predict. We show that this rule has a systematic blind spot: positions whose type is predictable but whose value must be computed, such as the number after "=" in a worked math solution. Under tight patch budgets, entropy-triggered layouts skip these positions, and accuracy on them collapses. In Meta's BLT-1B with patch starts on 10% of bytes, the entropy rule puts a patch start at 16% of the computed results in GSM8K solutions and gets 19.0% of them exactly right; a boundary after each "=" at the same patch count gets 51.8%, and entropy combined with a label-free boundary-dependence signal gets 67.1% (default layout at 26% of bytes: 76.8%). The gap survives adapting BLT-1B to the budget with low-rank fine-tuning (32.9% vs 72.7%, three runs per rule, paired p < 1e-200) and grows with model size in byte models trained from scratch at a 10% budget: at 1M, 12M and 50M parameters, boundary dependence beats entropy on final answers by -1.6, +10.1 and +19.8 points, and at 50M it gets 35.9% of computed results against 13.9% (3 seeds each). BLT's entropy-jump rule helps neither target at 50M. The entropy trigger of Scratchpad Patching is likewise indistinguishable from random scratchpads on final answers (5.6% vs 5.9%, 5 seeds), while answer-start scratchpads give 38.1%. The effect is specific to computed values: copies and lookups gain little, and values the model cannot compute gain nothing. Boundary dependence, the rise in the model's own loss when a patch start is removed, measured per two-byte context, finds these positions without labels: combined with entropy it beats the hand-written rule on computed results.
Comments13 pages, 4 figures, 6 tables. Code, logs and results: https://github.com/nicoveraz/segresearch (archived: doi:10.5281/zenodo.23238215)