发表机构
German Research Center for AI (DFKI); Saarland University; Center for European Research in Trusted AI (CERTAIN)(德国人工智能研究中心; 萨尔大学; 欧洲可信人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出分词化应被视为输出监督,通过解耦输入与输出分词化的实验验证其影响,调查120篇CL论文发现多数未报告数值分词化,为分词化影响模型性能提供了原则性解释。
AI 中文摘要
语言模型中的分词化默认被视为一种输入预处理决策。我们认为这一框架并不完整:在自回归模型中,分词器的粒度决定了模型在单次前向传播中必须解决的内容,因此也决定了模型接收到的监督信号。这既会影响学习问题的难度,也会影响模型内部涌现的表征。我们通过一项受控实验对数值推理任务进行了验证,该实验采用了输入与输出分词化的新型解耦设计。正如输出监督视角所预测的,任务性能、训练动态以及模型内部状态的差异均由输出分词化诱导,且在很大程度上不受输入分词化的影响。这在实践中可能具有重要意义,因为采用不同分词化策略的模型不仅在输入表征上存在差异,其训练时的任务也有所不同。因此,模型间的比较可能部分反映了任务定义而非模型能力。对120篇近期关于数值推理的CL论文的调查证实,这一情况很少被提及:仅约10%的论文报告了所评估模型的数值分词化方式,而69%的论文在跨分词化(即监督) regime进行比较时未报告相关信息。已有研究证实分词化会持续影响模型性能,但尚未有原则性的解释说明其原因。我们认为,将分词化视为输出监督的框架提供了这一解释。
英文摘要
Tokenization in language models is treated by default as an input preprocessing decision. We argue that this framing is incomplete: in autoregressive models, tokenizer granularity determines what the model must resolve in a single forward pass, and therefore the supervision signal it receives. This affects both the difficulty of the learning problem and the representations that emerge inside the model. We test this in a controlled experiment on numeric reasoning with a novel decoupling of input and output tokenization. As the output supervision view predicts, differences in task performance, training dynamics, and model internals are induced by output tokenization and largely invariant to input tokenization. This may matter in practice, because models with different tokenization strategies differ not only in input representation but in the task they were trained on. Comparisons between models may thus partly reflect task definition rather than ability. A survey of 120 recent *CL papers on numeric reasoning confirms that this is rarely acknowledged: only about 10% report the numeric tokenization of the models they evaluate, while 69% compare across tokenization, and thus supervision, regimes without reporting it. While prior work documents that tokenization consistently affects model performance, there is no principled account of why. We argue that framing tokenization as output supervision provides that account.
CommentsAccepted to EMNLP 2026 Main Conference