arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05597cs.CL

超越词面:面向高效语言模型的组合式分词

More Than Words: Compositional Tokenization for Efficient Language Models

发表机构耶路撒冷希伯来大学
查看机构详情
  • The Hebrew University of Jerusalem(耶路撒冷希伯来大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuval Reif, Guy Kaplan, Roy Schwartz

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出组合式分词方法CoBPE,将短语表示为基础词元加修饰符,在同等计算下缩短序列30%并提升下游性能1.2个百分点,为高效语言模型开辟新设计空间。

中文摘要 AI 辅助

语言模型以词元为单位顺序处理并生成文本,分词器决定了每个推理步骤覆盖多少文本。在标准分词下,诸如“On the table.”这样的短英文短语通常被生成为四个独立的预测,分别对应介词(On)、冠词(the)、名词(table)和标点(.),其中每个预测消耗一个序列位置并增加推理成本。我们提出了CoBPE,一种组合式分词方法,它将此类短语表示为一个词汇基础词元(table),并附加一组可复用的表面修饰符,在输入时于嵌入空间中进行组合,在输出时联合预测。在780M和1.3B规模下从头开始的受控预训练中,与标准BPE相比,在匹配的训练计算量下,CoBPE将序列缩短了30%,并将平均下游性能提高了1.2个百分点。我们的结果表明,当前通过词元序列表达的部分内容可以转而通过结构化表示来建模,从而为更高效、能力更强的语言模型开辟了广阔的设计空间。

英文摘要

Language models process and generate text sequentially in token units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as "On the table." is usually produced as four separate predictions for the preposition (On), article (the), noun (table), and punctuation (.), where each consumes a sequence position and adds inference cost. We introduce CoBPE, a compositional tokenization approach that represents such phrases as a lexical base token (table) attached with a small set of reusable surface modifiers, composed in embedding space at input and predicted jointly at output. In controlled pretraining from scratch at 780M and 1.3B scales, CoBPE shortens sequences by 30% and improves average downstream performance by 1.2 points relative to standard BPE under matched training compute. Our results suggest that part of what is now expressed through token sequences can instead be modeled through structured representations, opening a broad design space for more token-efficient and capable language models.

补充信息

↑