语言模型评估器在不同语言间存在偏差
Lower-Resource, Higher Scores: Language Bias in LLM Evaluators
浏览论文内容
中文总结 AI 辅助
研究发现语言模型评估器在多语言环境中存在偏差,不同语言评分差异显著,与语言资源水平相关,成对准确率无法检测到这些偏差,还探究了资源少的语言得分高的原因,揭示了语言层面的结构错位。
中文摘要 AI 辅助
语言模型评估器(经过训练的奖励模型和作为评判的提示语言模型)通常通过成对准确率进行验证。在多语言环境中,这基于高成对准确率意味着可靠、语言中立评分的前提。但研究表明该假设不成立。通过对23种语言的语义相同的指令 - 响应对进行实验,发现多语言评估器对不同评估语言给出显著不同分数。偏差具有统计学意义且在不同架构和训练范式的八个开放权重评估器中一致,在前沿评判中也存在,且与语言资源水平强烈相关,资源少的语言评分更宽松。这些偏差成对准确率检测不到,评估器成对准确率超90%,但在全局决策阈值下不同语言接受率差异达43%。研究还探讨了资源少的语言得分高的原因,发现模型不确定性与之有关,且偏差是结构上的语言层面错位,不能仅由内容难度解释。
英文摘要
LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwise accuracy. In a multilingual setting, this operates under the premise that high pairwise accuracy implies reliable, language-neutral scoring. We show that this assumption does not hold. We conduct experiments with semantically identical instruction-response pairs across 23 languages, and find that multilingual evaluators assign significantly different scores to different evaluation languages. The bias is statistically significant and consistent across eight open-weight evaluators of different architectures and training paradigms, persists in frontier judges, and is strongly correlated with language resource level: lower-resource languages are scored more generously. Meanwhile, these biases are invisible to pairwise accuracy: evaluators achieve above 90% pairwise accuracy, yet have up to 43% difference in acceptance rate across languages under a global decision threshold, meaning, for instance, that harmful content in lower-resource languages is more likely to pass safety filters. Per-language thresholds would require language identification, which can be defeated by code-switched prompts. We then investigate why lower-resource languages receive higher rather than lower scores, and we find that model uncertainty is linked with the effect: models tend to give higher scores when less confident, both under negative log-likelihood and under token-free uncertainty measures; however, language identity remains a significant predictor after controlling for uncertainty, and the bias cannot be explained away by content difficulty alone, but is a structural, language-level misalignment.
发表机构
- University of Cambridge(剑桥大学)
- Language Technology Lab(语言技术实验室)
机构由 AI 辅助整理,请以论文原文为准。