AI 中文总结
本文揭示基于查询复杂度的LLM路由存在语域偏差,非标准英语因省略功能词显得更短而被路由到低能力模型,且各层级模型本身对非标准语域回答准确度更低,加剧了不公平。
AI 中文摘要
大语言模型服务日益增多地使用一种廉价的查询复杂度估计,将每个查询路由到多个能力不同的模型之一,将简单查询发送给小模型,将困难查询发送给大模型。我表明这一路由步骤并非语域中立的:以非标准英语语域(非裔美国英语或第二语言写作者的英语)撰写的文本,被系统性地分配到比语义等价的标准英语版本查询更低的能力层级。该效应由一种特定且常见的路由信号——输入长度——驱动,因为非标准语域省略了功能词,从而看起来更短、因此更简单;其他复杂度信号不携带此效应。我在37,704个真实学习者句子对和一个受控平行语料库上证明了这种差异。然后,我在设备、边缘和云模型阶梯上衡量了质量后果,发现伤害由普遍的模型偏差驱动:每一层级,包括前沿云模型,对非标准语域查询的回答准确度显著更低,而在此基准上,路由决策本身的边际质量成本并不显著。因此,基于复杂度的路由加剧了模型本已服务最差的用户的暴露程度。
英文摘要
Large language model services increasingly route each query to one of several models of differing capability, using a cheap estimate of query complexity to send easy queries to small models and hard queries to large ones. I show that this routing step is not register neutral: text written in a non-standard English register, African American English or the English of second-language writers, is systematically assigned a lower-capacity tier than a meaning-equivalent standard-English version of the same query. The effect is driven by a specific, common routing signal, input length, because non-standard registers omit function words and thus look shorter and therefore simpler; other complexity signals do not carry it. I demonstrate the disparity on 37,704 authentic learner sentence pairs and on a controlled parallel corpus. I then measure the quality consequence on a device, edge, and cloud model ladder and find that the harm is driven by pervasive model bias, every tier, including a frontier cloud model, answers non-standard-register queries significantly less accurately, while the marginal quality cost of the routing decision itself is not significant on this benchmark. Complexity-based routing thus compounds the exposure of the users that the models already serve worst.
Comments4 pages, 2 figures. Code: https://github.com/SimranKoul2026/register-bias-llm-routing