语言编码的速率-效用前沿:在受控语言内容下比较词元、字节和像素
Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content
浏览论文内容
中文总结 AI 辅助
研究在内容和下游能力受控时语言编码保留了什么,利用多语言平行句子通过共享瓶颈比较词元、字节和像素以追踪速率-效用前沿,评估三种效用,发现各编码在不同方面表现不同,选择编码需考虑多因素进行速率-效用权衡。
中文摘要 AI 辅助
语言模型将文本编码为子词元、原始字节或渲染像素,但这些编码通常在建模约束下进行比较,而这些约束会使不同语言的模型接触到不同数量的语言内容。本文提出在内容和下游能力均受控制的情况下,研究每种编码保留了什么。利用13种语言和5种脚本的经过验证的平行句子,通过一个共享瓶颈比较词元、字节和像素,该瓶颈的宽度被扫描以追踪速率-效用前沿。这区分了三个常被混淆的量:编码创建的输入位置数量、编码后可用的潜在容量以及压缩后留存的与任务相关的信息。评估了三种效用:表面形式保留、跨语言句子对齐和主题分类。没有一种编码在所有任务或容量范围内占主导地位。像素最能保留表面形式,字节最能保留跨语言对齐,尤其是在同脚本多语言设置中,词元最支持主题预测。这些性能不能仅由序列长度来解释。短输入可能会丢弃有用的含义,而长输入可能会保留易于压缩的信息。因此,选择一种编码不是对词元、字节或像素的固定偏好,而是一种取决于任务、语言组合、容量范围和计算预算的速率-效用权衡。
英文摘要
Language models encode text as subword tokens, raw bytes, or rendered pixels, but these encodings are usually compared under modeling constraints that expose different amounts of linguistic content to models across different languages. We instead ask what each encoding preserves when both the content and the downstream capacity are controlled. Using verified parallel sentences across thirteen languages and five scripts, we compare tokens, bytes, and pixels through a shared bottleneck whose width is swept to trace rate-utility frontiers. This separates three quantities that are often conflated: the number of input positions an encoding creates, the latent capacity available after encoding, and the task-relevant information that survives compression. We evaluate three utilities: surface form preservation, cross-lingual sentence alignment, and topic classification. No encoding dominates across tasks or capacity regimes. Pixels preserve surface form best, bytes preserve cross-lingual alignment best, especially in same-script multilingual settings, and tokens support topic prediction best. These performances are not explained by sequence length alone. Short inputs can discard useful meaning, while long inputs can preserve information that compresses well. Choosing an encoding is therefore not a fixed preference for tokens, bytes, or pixels, but a rate-utility tradeoff that depends on the task, language mix, capacity regime, and compute budget.
发表机构
- Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。