arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15454cs.AI

基于分层语言模型的动态多字节预测

Dynamic Multi-Byte Prediction With Hierarchical Language Models

Abraham Toluwase Owodunni, Chibuzor Okocha, Christan Grant, Tomasz Limisiewicz, Sachin Kumar

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对分层语言模型逐字节生成推理速度慢的问题,提出多字节预测方法,通过可变长度预测窗口与因果注意力掩码实现并行字节预测,在多任务中达成性能与吞吐量的最优权衡。

中文摘要 AI 辅助

字节级分层语言模型(LMs)近年来已成为流行的子词分词模型的稳健替代方案,但逐字节生成仍是推理速度的瓶颈。为解决该问题,本文提出多字节预测(MBP),可并行生成多个字节,在几乎不影响性能、不增加额外参数的前提下加快推理速度。MBP基于流行的多令牌预测(MTP)范式,有两项关键创新:一是引入与分层语言模型的潜在令牌(即片段)对齐的可变长度预测窗口;二是实现新颖的注意力掩码方案,可在不违反因果性的前提下实现并行字节预测。实验在生成式任务、指令遵循、问答、摘要和机器翻译中验证,MBP在性能与推理吞吐量间实现了帕累托最优权衡,达成最优的性能-吞吐量平衡。

英文摘要

Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword tokenization. However, generating one byte at a time remains a bottleneck for inference speed. To address this, we introduce multi-byte prediction (MBP), which generates multiple bytes in parallel, speeding up inference with minimal performance impact and no additional parameters. MBP builds on the popular multi-token prediction (MTP) paradigm with two crucial innovations. First, we introduce a variable-length prediction window that aligns with the latent tokens, or segments, of a hierarchical LM. Second, we implement a novel attention-masking scheme that enables parallel byte prediction without violating causality. We show that multi-byte prediction strikes a Pareto-optimal trade-off across multiple generative tasks, instruction following, question answering, summarization, and machine translation, achieving the best trade-off between performance and inference throughput.

发表机构

  • The Ohio State University(俄亥俄州立大学)
  • University of Florida(佛罗里达大学)
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑