arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更少的词,而非更少的 token:衡量梵语每个命题的 token 化惩罚

Fewer Words, Not Fewer Tokens: Measuring the Sanskrit Tokenization Penalty per Proposition

Devansh Sharma

arXiv 2609.12960首次发表:更新:

AI 中文总结

本研究衡量梵语在子词 token 化下的效率惩罚,发现已部署 token 化器下每个命题的 token 成本高于英语,但匹配训练后差距缩小,主要源于字符长度差异。

AI 中文摘要

梵语将格、数、人称和时态融合到词尾中,并将子句链接成复合词,因此每个词的信息密度很高。这种密度能否在子词 token 化中幸存下来,是一个单独的问题,需要按意义单位而非按词来提出。在相同的 FLORES-200 devtest 内容上,在使用词汇量为 200,019 个或更多 id 的已部署 token 化器下,梵语的 token 数量是英语的 1.774-2.187 倍,但仅为印地语的 1.325-1.353 倍。相对于已部署的英语 token 化器,在当代散文中,基于梵语训练的 BPE 分支每个命题的成本看起来比英语便宜(0.887)。与匹配的英语对照(即在同一语料库的英语一侧训练的相同算法和词汇)相比,这种翻转消失了:在 32,000 和 64,000 个片段下,所有 8 个匹配对(每个大小匹配的分支都同时与配对匹配和字节匹配的对照进行比较)在散文上均高于 1.0,且 95% 置信区间排除该值。随着词汇量的增长,差距缩小:在 128,000 个片段时,BPE 对在领域内读数为 0.983,而在领域外(1.025)和 FLORES(1.116)上仍高于均等水平。该比率可分解为字符长度比率和每字符 token 比率,后者在整个过程中接近 1:在匹配的 token 化下幸存下来的是字符级长度,而梵语散文在 SLP1 中相对于英语缺乏这种长度(1.028),而梵语诗歌则具有(0.596)。稳健的陈述是关于已部署实践的:在当代散文和 FLORES 上,梵语一侧使用 SLP1 与已部署的 o200k 英语枢轴相比,在人们实际使用的 token 化器下,梵语每个命题花费 1.831-2.899 个英语 token。代码、结果快照和这里的每个表格都是公开的。

英文摘要

Sanskrit fuses case, number, person and tense into word endings and chains clauses into compounds, so it is information-dense per word. Whether that density survives subword tokenization is a separate question, to be asked per unit of meaning rather than per word. On identical FLORES-200 devtest content, Sanskrit costs 1.774-2.187 times the English tokens under deployed tokenizers with vocabularies of 200,019 ids or more, but only 1.325-1.353 times the Hindi tokens. Against a deployed English tokenizer, Sanskrit-trained BPE arms then look cheaper per proposition than English on contemporary prose (0.887). Against a matched English control, the same algorithm and vocabulary trained on the English side of the same corpus, that flip disappears: at 32,000 and 64,000 pieces all 8 matched pairs, each size-matched arm against both a pair-matched and a byte-matched control, sit above 1.0 on prose with 95% intervals excluding it. The gap closes as the vocabulary grows: at 128,000 pieces the BPE pair reads 0.983 in domain while staying above parity out of domain (1.025) and on FLORES (1.116). The ratio factorises into a character-length ratio and a tokens-per-character ratio, the second near 1 throughout: what survives matched tokenization is character-level length, which Sanskrit prose lacks over English in SLP1 (1.028) and Sanskrit verse has (0.596). The robust statement is about deployed practice: on contemporary prose and on FLORES, with the Sanskrit side in SLP1 against the deployed o200k English pivot, Sanskrit costs 1.831-2.899 English tokens per proposition under the tokenizers people actually ship. Code, the results snapshot and every table here are public.

Comments20 pages, of which 8 are the body; 4 figures, 18 tables. Code, data pipeline and the full results snapshot: https://github.com/DS436/sanskrit-token

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑