AI 中文总结
EntropyMoE针对无分词器字节级LLM的块计算局限性,提出基于块熵的稀疏专家路由MoE架构,在降低位每字节的同时保持下游准确率,扩展了MoE建模的应用范围。
AI 中文摘要
近期的字节级大语言模型(LLM)通过将字节分组为动态大小的块,使得无分词器建模的竞争力不断提升。然而,现有的字节块架构仍对每个块应用相同的密集前馈计算,这种统一计算无法使模型容量适配块语义和粒度的变化。我们针对该局限性提出了EntropyMoE,一种专为动态字节块设计的混合专家(MoE)架构。EntropyMoE将全局块Transformer中的密集前馈模块替换为Top-K专家层,每个动态块作为专家路由的基本单元,其字节覆盖度决定了对工作量统计的贡献;路由器直接基于块熵选择专家,利用构成动态块的相同粒度信号组织稀疏计算,块熵和长度共同定义了调节专家专业化的特征空间。实验表明,在匹配的密集和稀疏基线中,EntropyMoE实现了最低的保留位每字节(bits-per-byte),同时保持了相当的下游准确率。这些结果确立了块熵作为稀疏条件计算的有效路由坐标的有效性,并将混合专家建模扩展到无分词器表示之外。
英文摘要
Recent byte-level large language models (LLMs) have made tokenizer-free modeling increasingly competitive by grouping bytes into dynamically sized patches. However, existing byte-patch architectures still apply the same dense feed-forward computation to every patch. This uniform computation cannot adapt model capacity to variations in patch semantics and granularity. We address this limitation with EntropyMoE, a Mixture-of-Experts (MoE) architecture designed for dynamic byte patches. EntropyMoE replaces the dense feed-forward modules in the global patch Transformer with Top-K expert layers. Each dynamic patch serves as the basic unit of expert routing, and its byte coverage determines its contribution to workload accounting. The router selects experts directly from patch entropy, using the same granularity signal that underlies dynamic patch construction to organize sparse computation. Patch entropy and length jointly define the feature space for regulating expert specialization. Experiments show that EntropyMoE achieves the lowest held-out bits-per-byte among matched dense and sparse baselines while maintaining comparable downstream accuracy. These results establish patch entropy as an effective routing coordinate for sparse conditional computation and extend Mixture-of-Experts modeling beyond tokenizer-based representations.