AI 中文总结
该研究提出字节级语言模型可解耦下一字节分布与边界位置分布,设计实验验证该假设,主张采用字节级接口作为共享标准以实现模型间低成本的能力传递与边界重塑。
AI 中文摘要
通常,字节级语言模型的优势被认为在于鲁棒性、多语言公平性以及字符级技能。我们指出其另一种结构性优势:由于它们读写字节,任意两个字节级语言模型共享同一输出空间,因此它们之间的知识传递是精确的,且不依赖于各自原始的分词方式。我们假设字节级模型生成的两种分布——一种针对下一个字节,另一种针对其片段边界的位置——可以被解耦,并且几乎可以独立更改。模型可以吸收教师模型的能力,同时保留自身的边界;或者改变边界的放置方式,同时保留自身的能力。我们设计了两项可验证该假设的实验,并附带对实验所基于属性的初步测量。我们认为,学界应转向字节级接口作为共享标准:若该假设成立,一旦字节级模型成为主流,模型间的能力传递与边界重塑将变得廉价且常规,无需再受当前阻碍的各模型专属分词器的限制。
英文摘要
Byte-level language models are usually argued for on the grounds of robustness, multilingual fairness, and character-level skills. We point to a different, structural advantage: because they read and write bytes, any two of them share an output space, so knowledge transfer between them is exact and independent of how either was originally tokenized. We hypothesize that the two distributions a byte-level model produces, one over the next byte, one over where its patch boundaries fall, can be disentangled and changed almost independently. A model could absorb a teacher's capability while keeping its own boundaries, or change how it places those boundaries while keeping its capabilities. We lay out the two experiments that would settle the hypothesis, alongside preliminary measurements of the properties they rest on. We argue that the community should move toward a byte-level interface as a shared standard: if the hypothesis holds, then once byte-level models are the norm, transferring capabilities and reshaping boundaries between them become cheap and routine, free of the per-model tokenizer that blocks them today.