arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生命周期最优分词:词汇表大小作为与部署 regime 相关的基础设施参数

Lifecycle-Optimal Tokenization: Vocabulary Size as a Deployment-Regime-Dependent Infrastructure Parameter

Rima Mittal, Ankit Gubrani, Satyanarayana Kakollu

arXiv 2608.11361首次发表:更新:

AI 中文总结

该研究发现 LLM 分词器的最优词汇表大小随部署 regime 变化,提出生命周期部署成本模型,经实验证实其与训练最优值差异最高达 16 倍,且最优范围内质量损失极小,为 LLM 部署提供词汇表容量规划指导。

AI 中文摘要

分词器词汇表大小是大语言模型(LLM)基础设施中的一项基础设计选择,但通常基于惯例在训练时固定,而非通过部署分析确定。本文表明,成本最优词汇表并非常数,而是 serving regime 的函数。我们将总部署成本形式化为 $C_{lifecycle}(V) = C_{train}(V) + \lambda \cdot C_{infer}(V, B)$,其中 $\lambda$ 为推理量,$B$ 为 serving 批大小。通过在两类 GPU 上开展受控实验,覆盖内存受限到计算受限 regime(A10G 的 ridge ≈ 117 FLOP/byte;A100 的 ridge ≈ 183 FLOP/byte),我们证实:(1)推理最优词汇表随 serving 批大小变化 16 倍,当 $B=1$ 时为 32k,当 $B=64+$ 时为 524k,这由 $V \times d$ 未嵌入矩阵读取的摊销驱动;(2)在 13 亿至 23 亿参数的模型规模下,质量(每字节比特数,BPB)在 $V=65$k 时达到最优,证实了规模依赖的词汇表偏好;(3)对于生产部署,生命周期最优词汇表与训练最优词汇表的差异最高达 16 倍。在最优范围内,质量近似不变(BPB 波动 <2%),使得词汇表成为纯系统优化,在所测范围内无质量损失。我们的结果提供了可操作的容量规划指导:设备端部署($B=1$)应使用 $V≈32$k;数据中心 serving($B≥64$,$λ≥10$)应使用 $V≈131$-262k。

英文摘要

Tokenizer vocabulary size is a foundational design choice in large language model (LLM) infrastructure, yet it is typically fixed at training time based on convention rather than deployment analysis. We show that the cost-optimal vocabulary is not a constant but a function of the serving regime. We formalize total deployment cost as $C_{lifecycle}(V) = C_{train}(V) + λ\cdot C_{infer}(V, B)$, where $λ$ is inference volume and $B$ is the serving batch size. Through controlled experiments on two GPU families spanning the memory-bound to compute-bound regimes (A10G, ridge $\approx$ 117 FLOP/byte; A100, ridge $\approx$ 183 FLOP/byte), we demonstrate: (1) the inference-optimal vocabulary shifts 16x with serving batch, from 32k at $B=1$ to 524k at $B=64+$, driven by amortization of the $V \times d$ unembedding matrix read; (2) at 1.3-2.3B model scale, quality (bits per byte, BPB) is optimized at $V=65$k, confirming scale-dependent vocabulary preference; (3) the lifecycle-optimal vocabulary diverges from training-optimal by up to 16x for production deployments. Quality is approximately invariant across the optimal range ($<$2% BPB spread), making vocabulary a pure systems optimization with no quality penalty in the measured range. Our results provide actionable capacity planning guidance: on-device deployments ($B=1$) should use $V \approx 32$k; datacenter serving ($B \geq 64$, $λ\geq 10$) should use $V \approx 131$-262k.

Comments6 pages, 3 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑