arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自动补全分词器的试点研究

A Pilot Study of Autocompleting Tokenizers

Samuel Wexler, Mark Hopkins

arXiv 2608.15080首次发表:更新:

发表机构

Williams College(威廉姆斯学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究受自动补全书写系统启发,提出用轻量级自回归字节语言模型压缩Transformer输入的方案,在多语言机器翻译任务上验证其可减少序列长度且不降低翻译质量,适用于不同书写系统。

AI 中文摘要

现代输入方法通常依赖自动补全来省略可从局部上下文恢复的信息。受这些自动补全辅助书写系统的启发,我们研究是否可以以类似方式压缩Transformer输入。字节级分词是子词分词的一种简单、语言无关的替代方案,但其较长的输入序列通常会导致计算成本增加和模型质量下降。我们提出一种压缩方案,该方案采用轻量级自回归字节语言模型,在Transformer处理前识别并移除可从周围上下文轻松预测的字节。得到的压缩表示随后被提供给标准编码器-解码器Transformer作为输入。机器翻译实验表明,在不降低翻译质量的情况下,可以省略相当一部分源语言字节。在英法翻译任务中,我们的最佳方法在保持翻译性能的同时,将源序列长度减少了近三分之一。在芬英、俄英和中英翻译任务上的额外实验表明,该方法可在不同书写系统和形态类型学上泛化,在0.47至0.67的压缩率下产生相当或更好的翻译质量。这些发现表明,许多输入字节具有足够的可预测性,可以隐式表示而非显式表示,为减少与字节级模型相关的序列长度开销提供了一种简单机制。

英文摘要

Modern input methods routinely rely on autocomplete to omit information that can be recovered from local context. Inspired by these autocomplete-assisted writing systems, we investigate whether Transformer inputs can be compressed in a similar manner. Byte-level tokenization offers a simple and language-independent alternative to subword tokenization, but its longer input sequences typically result in increased computational cost and reduced model quality. We propose a compression scheme that employs a lightweight autoregressive byte language model to identify and remove bytes that are easily predictable from their surrounding context before Transformer processing. The resulting compressed representation is then provided as input to a standard encoder--decoder Transformer. Experiments on machine translation show that a substantial fraction of source-language bytes can be omitted without degrading translation quality. On English--French, our best method preserves translation performance while reducing source sequence length by nearly one-third. Additional experiments on Finnish--English, Russian--English, and Chinese--English demonstrate that the approach generalizes across diverse writing systems and morphological typologies, yielding comparable or improved translation quality at compression ratios between 0.47 and 0.67. These findings suggest that many input bytes are predictable enough to be represented implicitly rather than explicitly, providing a simple mechanism for reducing the sequence-length overhead associated with byte-level models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑