发表机构
University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在自然语言处理全流程追踪“方言税”,发现现代语言模型在分词、预训练等各环节均存在方言偏见,且字符级分词器无法消除相关差距,方言税由流程各环节共同累积而成。
AI 中文摘要
语言模型(LMs)中系统性的方言性能差距已得到充分记录,但现代语言建模流程中这些差异的来源仍不清楚。本研究在自然语言处理流程中追踪这种“方言税”。使用保持意义固定但表层形式变化的平行英语方言语料库,我们首先确认语言模型将匹配的标准美式英语(SAE)和方言文本识别为语义等价。然而,我们发现对应下游性能差距的进一步表征差距:在所有模型家族和代际中,现代语言模型在分词、预训练、后训练和推理过程中仍不平等地编码方言文本。值得注意的是,通过字符级反事实分词器绕过传统子词分词,既未消除输入和输出不对称,也未消除方言准确率差距。预训练期间,方言对相比完全不相关的SAE文档对会引发更发散的梯度更新,表明模型从语义等价的方言内容中学习比从不相关的SAE文档中学习更困难。后训练期间,奖励模型表现出语境依赖、不稳定的方言偏好:为孤立的非裔美国人方言英语(AAVE)专属标记分配比SAE专属标记更高的价值,而完整推理语境则会受到任务和模型依赖的方言惩罚。总体而言,我们的发现表明,方言税并非由任何单一步骤单独编码和积累,而是在语言建模过程的每一步都存在。
英文摘要
Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear. Our study traces this "dialect tax" across the natural language processing pipeline. Using parallel English dialect corpora that hold meaning fixed while varying surface form, we first confirm that LMs recognize matched Standard American English (SAE) and dialectal texts as semantically equivalent. However, we discover further representational gaps corresponding to downstream performance gaps. Across model families and generations, modern LMs still encode dialectal texts unequally during tokenization, pre-training, post-training, and inference. Strikingly, bypassing traditional subword segmentation via a character-level counterfactual tokenizer removes neither input and output asymmetries nor dialectal accuracy gaps. During pre-training, dialect pairs induce more divergent gradient updates than pairs of entirely unrelated SAE documents, indicating that models find semantically equivalent dialectal content harder to learn from than unrelated SAE documents. During post-training, reward models show contextual, unstable dialect preferences, assigning higher values to isolated AAVE-exclusive tokens than to SAE-exclusive tokens, while full reasoning contexts receive task- and model-dependent dialect penalties. Overall, our findings suggest that the dialect tax is encoded and accumulated not by any one step in isolation, but at every step of the language modeling process.
CommentsTo be published at EMNLP 2026 under the full author name "Elle"