发表机构
The Chinese University of Hong Kong, Shenzhen; Shenzhen Loop Area Institute(香港中文大学(深圳); 深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究量化预分词边界带来的压缩代价,通过上下界与证书方法,发现边界使最优token数增加28.3%-36.8%,并引入边界许可证研究中间策略,同时区分压缩与预测的影响。
AI 中文摘要
预分词(pre-tokenisation)限制了哪些文本片段能够成为预测单元,但当分词器仅在相同边界下进行比较时,其压缩成本被掩盖了。我们通过从两个方向界定最小 token 数量来衡量这一成本,分别在有和没有正则表达式边界规则的情况下进行。token 出现位置上的非负价格通过最短路径和词汇预算选择产生一个下界;对所有价格取最大值可恢复线性规划松弛,而一个独立的整数检查器则对报告的值进行认证。在英文维基百科上,边界使最优 token 数量增加了 28.3% 至 36.8%。字节对编码(Byte Pair Encoding)比受约束的下界高 2.1%,但比无约束的下界高 10.9%。压缩和预测偏好不同的词典:在 85M 非嵌入参数和匹配的训练 token 预算下,无约束拟合在共同的无约束解码器下,在配对研究的所有 12 种语言中产生更高的平均保留每字节比特数(held-out bits per byte),在独立调优和评估下则为 12 种语言中的 11 种。为了研究中间边界策略,我们引入了边界许可证(boundary licences),它限制允许跨越切口的词汇条目,并允许相同形式的证书。在独立的英文和中文拟合语料上,许可 10% 的词汇预算可恢复移除所有切口所实现的 token 数量减少的 85.2% 和 100.0%。这些结果量化了边界的压缩成本,同时将其与所得 token 单元的预测质量分开。
英文摘要
Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported values. On English Wikipedia, boundaries increase the optimal token count by 28.3--36.8\%. Byte pair encoding lies 2.1\% above the constrained lower bound, but 10.9\% above the unrestricted bound. Compression and prediction favour different dictionaries: at 85M non-embedding parameters and matched training-token budgets, unrestricted fitting yields higher mean held-out bits per byte under a common unrestricted decoder in all 12 languages in the paired study and 11 of 12 under independent tuning and evaluation. To study intermediate boundary policies, we introduce boundary licences, which limit the vocabulary entries permitted to cross cuts and admit the same form of certificate. On separate English and Chinese fitting corpora, licensing 10\% of the vocabulary budget recovers 85.2\% and 100.0\% of the achieved token-count reduction from removing all cuts. These results quantify the compression cost of boundaries while separating it from the prediction quality of the resulting token units.