arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

字节语言模型:扩展、涌现抽象与信息分配

Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation

Jie Wang, Shiwei Luo, Qi Zhang, Yuanbin Wu

arXiv 2610.05978首次发表:更新:

发表机构

East China Normal University; Fudan University(华东师范大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究字节级无分词器Transformer,发现其通过令牌叠加和哈希嵌入在扩展时优于子词模型,并自发形成局部文本抽象,利用这些结构可提升投机解码效率。

AI 中文摘要

无分词器语言模型通过直接以字节形式建模文本,消除了固定分词器的归纳偏置,但由此产生的更长序列大幅增加了计算量,并消除了显式的文本抽象。我们探究这些额外计算是否具有实用性,以及标准Transformer能否学习到分词所提供的抽象。我们在没有专门分词相关架构的Transformer上研究这些问题。通过令牌叠加训练和哈希嵌入,字节Transformer在模型规模扩展时持续优于子词Transformer。我们进一步发现,字节Transformer构建了类似外部分词器的局部文本抽象:一组类似分割的位置被用于收集局部上下文表示,且将多达25%的中间层限制为这些局部表示不会降低下游性能。最后,这些学习到的结构导致了高度非均匀的生成难度,不确定性集中在局部结构边界附近;利用这些结构进行投机解码,比子词Transformer多接受3.4倍的令牌。

英文摘要

Tokenizer-free language models remove the inductive bias of fixed tokenizers by modeling text directly as bytes, but the resulting longer sequences substantially increase computation and eliminate explicit text abstractions. We ask whether this additional computation can be useful, and whether standard Transformers can learn the abstractions that tokenization provides. We study these questions on Transformers without specialized tokenization-related architectures. With token-superposition training and hash embeddings, byte Transformers consistently outperform subword Transformers as model size scales. We further find that byte Transformers build local text abstractions as external tokenizers: a set of segmentation-like positions are used to collect local context representations, and restricting up to $25\%$ of intermediate layers to these local representations preserves downstream performance. Finally, these learned structures induce highly non-uniform generation difficulty, with uncertainty concentrated near local structure boundaries; exploiting them for speculative decoding yields $3.4\times$ more accepted tokens than in subword Transformers.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑