发表机构
EPFL(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出用显式词边界标记替代标准空格约定,缓解了子词词汇的条目重复问题,虽未提升分词压缩率,但可改善语言建模性能。
AI 中文摘要
在使用空格的书写系统中,子词分词器会将许多常见词表示两次:一次带前置空格,一次不带。模型中这两个条目有独立的嵌入,因此同一词的出现会被划分到独立训练的行中,甚至二者对字符串的分词方式可能不同:“ together”可能是单个条目,而不带前置空格的同一词会被分词为“to|gether”。大小写会进一步将词划分为多达六种形式。我们提出了一种替代标准空格约定的方案,使用显式词边界标记来防止此类重复。词由边界标记分隔,词之间的空格表示为此类标记对。两个移位标记分别用于标题大小写和大写,允许同一词的内部表示在不同场景下重复使用。采用该约定缓解了条目重复问题,但未提升分词压缩率:对于两种词汇学习算法,在六种语言上平均而言,最佳标记方案的字符数/词素数与基线的差距在1%以内。不过它确实提升了语言建模性能,所有测试的标记方案在下游任务中均达到比基线更低的每字节比特数,表明重复带来了压缩率未捕捉到的成本。
英文摘要
Subword tokenizers represent many common words twice in space-using writing systems, once with a leading space and once without. The two entries have separate embeddings in models, so occurrences of one word are divided across rows that are trained independently, and the two forms need not even segment the string the same way: " together" may be a single entry while the same word without a preceding space is tokenized as "to|gether". Capitalization divides a word further, into as many as six forms. We introduce an alternative to standard whitespace conventions using an explicit word boundary marker, which prevents such duplication. Words are delimited by the boundary markers, and spaces between words are represented as pairs of such markers. Two shift codes do the same for title case and upper case, allowing one internal representation of a word to be re-used across different settings. Switching to this convention mitigates the duplicate-entry issue, but does not improve tokenization compression: for both vocabulary-learning algorithms, the best marker scheme stays within one percent of the baseline in characters per token, averaged across six languages. It does result in better language modeling performance. Every marker scheme tested downstream reaches lower bits per byte than the baseline, suggesting that duplication carries a cost that compression does not capture.
CommentsCode available at https://github.com/sanderland/script_tok