Functionalizer:用于子词分词的无损函数分解
The Functionalizer: Lossless Functional Decomposition for Subword Tokenization
- Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Functionalizer通过将正字法和结构变体分解为可逆的操作码/操作数前缀流,实现无损子词分词,在减少词汇量需求的同时提升代码建模性能。
AI中文摘要:
标准子词分词器要么将单词的每种正字法变体(如hello、Hello、HELLO和Héllo)视为无关的词汇条目,这会使嵌入空间碎片化,要么通过有损归一化丢弃这些变体。我们提出了Functionalizer,一个无损预分词框架,在分词前将正字法和结构变体分解为组合的操作码/操作数前缀流:一个规范的基础词元(操作数)前缀以编码在Unicode私有使用区的参数化变换算子(操作码)。我们引入了涵盖大小写(CAPITALIZE)、变音符号(13个专用操作码)和字符重复(REPEAT、MULTIREPEAT)的算子,这些算子完全可逆。在六个自然语言和代码语料库上,Functionalizer在无约束条件下实现了完整的语料库覆盖,且词汇量显著更小,实际词汇槽需求最多减少16%。在序列长度方面,我们观察到明显的领域依赖权衡:它压缩了缩进密集的代码序列,但膨胀了自然语言散文序列。在25M参数的GPT-2规模模型上的初步下游评估表明,在此规模下,Functionalizer大幅提高了代码语法有效性,并改善了代码字符困惑度,同时在散文上保持了相似的文本连贯性。这些发现表明,函数分解可以成为词汇高效、结构感知语言建模的有效机制,并激励在生产规模上进行进一步验证。
英文摘要:
Standard subword tokenizers either treat every orthographic variation of a word (such as hello, Hello, HELLO, and Héllo) as unrelated vocabulary entries, which fragments the embedding space, or discard this variation through lossy normalization. We present the Functionalizer, a lossless pre-tokenizer framework that factors orthographic and structural variations into a compositional opcode/operand prefix stream before tokenization: a canonical base token (operand) prefixed by parametric transformation operators (opcodes) encoded in the Unicode Private Use Area. We introduce operators covering casing (CAPITALIZE), diacritics (13 dedicated opcodes), and character repetition (REPEAT, MULTIREPEAT), which are fully reversible. Across natural language and code corpora, the Functionalizer enables complete corpus coverage with significantly smaller vocabularies under unconstrained exhaustion conditions, reducing actual vocabulary slot requirements by up to 19.7%. Downstream evaluations on 98M-parameter GPT-2 models show that the Functionalizer improves Python code syntax validity (9.12% vs. 7.70%) while reducing duplicate n-gram repetition in natural language prose. These findings demonstrate that functional decomposition can be an effective mechanism for vocabulary-efficient, structurally aware language modeling, and motivate further validation at production scale.