打破字节级语言建模中的层级结构
Toppling the Hierarchy in Byte-level Language Modeling
浏览论文内容
中文总结 AI 辅助
本研究针对当前最优字节级语言模型的层级结构缺陷,通过消融实验发现纯字节级模型在字符操控任务上更优,明确了计算效率与细粒度字符理解的权衡关系。
中文摘要 AI 辅助
本研究探讨了近期字节级模型在字符完美操控方面存在的缺陷。当前最优的字节级模型采用层级结构,从字节级开始,下采样至词级,再上采样回字节。尽管这提升了训练与推理效率,但我们发现该层级设计本身限制了字符级理解,纯字节级模型在字符操控任务上始终优于其层级变体。将Transformer层拆解为注意力与前馈组件的消融实验进一步表明,字节级注意力是驱动这一现象的核心机制。综上,我们的结果解释了层级字节模型在字符级任务上的失败原因,并明确了计算效率与细粒度字符理解之间的权衡关系。
英文摘要
This work examines recent byte-level models and their failure to perfectly manipulate characters. State-of-the-art byte-level models use a hierarchical structure, starting at the byte level, downsampling to the word level, and then upsampling back to bytes. While this improves training and inference efficiency, we find that the hierarchical design itself limits character-level understanding, with pure byte-level models consistently outperforming hierarchical variants on character manipulation tasks. Ablating transformer layers into attention and feed-forward components further reveals that byte-level attention is the primary mechanism driving this behavior. Together, our results provide an explanation for the character-level failures of hierarchical byte models and establish a clear trade-off between computational efficiency and fine-grained character understanding.
发表机构
- School of Computation, Information and Technology, TU Munich(慕尼黑工业大学计算、信息与技术学院)
- Munich Center for Machine Learning(慕尼黑机器学习中心)
- Munich Data Science Institute(慕尼黑数据科学研究所)
机构由 AI 辅助整理,请以论文原文为准。