arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22432cs.CLcs.AI

多语言大语言模型评判器中的排名反转:一种无标签双中心化校准器

Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator

Alhasan Mahmood, Samir Abdaljalil, Hasan Kurban

首次发表
浏览论文内容

中文总结 AI 辅助

针对多语言LLM评判器因提示语言导致排名反转的问题,提出无标签共识校准法(CBC),可提升排名一致性及与人工偏好的契合度,验证了其下游实用性。

中文摘要 AI 辅助

多语言大语言模型(LLM)评判器会根据提示语言生成不同的评判器主干模型排名:在包含8种语言的“智能体即评判器(Agent-as-a-Judge)”基准测试中,排名最高的主干模型会随英语、阿拉伯语、汉语、印地语、日语、西班牙语、土耳其语和斯瓦希里语交替出现,且15组主干模型对中有7组表现出统计学意义上的成对排名反转。我们将此视为一种测量问题。多语言评判器得分可加性分解为任务难度、主干模型技能以及语言-主干模型交互项,其中交互项可通过对单元平均得分矩阵进行双中心化处理,无需人工标签即可恢复。我们明确提出了该估计器(共识校准法,Consensus-Based Calibration,CBC),给出了其有限样本收敛界,其方差常数为$(1-\frac{1}{m})(1-\frac{1}{k})$,且即使存在任务-语言交互,该估计器仍为无偏估计。在7920次评判运行(6个主干模型、8种语言、55个任务、3个框架)中,CBC将保留的跨任务排名一致性$\tau$从0.650提升至0.902,且在单语言决策中与保留的加性模型神谕完全一致,而原始方法的一致性仅为68.5%;这些均为一致性诊断,而非基于人工的正确性度量。在单独收集的M-RewardBench面板(7种语言、每种语言1500个样本、共10500个语言-样本实例、5个评判器)上,面板与公开人工黄金偏好的一致性从68.7%提升至76.6%(提升7.9个百分点,95%置信区间为[6.0,9.9]),这是其下游实用性的最强外部证据。该估计器是零和对比下标准双向方差分析的交互项恢复操作;我们的贡献在于将其应用为多语言LLM评判器的无标签事后校准器,给出明确的有限样本收敛界,以及即使在任务-语言设定错误下仍成立的无偏性结果。

英文摘要

Multilingual LLM judges produce different evaluator-backbone rankings depending on the prompt language: on an eight-language Agent-as-a-Judge benchmark, the top-ranked backbone alternates across English, Arabic, Chinese, Hindi, Japanese, Spanish, Turkish, and Swahili, and 7 of 15 backbone pairs show statistically significant pairwise rank reversal. We treat this as a measurement problem. The multilingual judge score decomposes additively into task difficulty, backbone skill, and a language-backbone interaction term, the last of which is recoverable without human labels by double-centering the cell-mean score matrix. We make this estimator (\textbf{Consensus-Based Calibration}, CBC) explicit, give an $O(1/\sqrt{n})$ finite-sample concentration bound with variance constant $(1-\tfrac{1}{m})(1-\tfrac{1}{k})$, and show that it is unbiased even when task-language interactions are present. Across 7{,}920 judge runs (6 backbones, 8 languages, 55 tasks, 3 frameworks), CBC raises held-out cross-task rank consistency $τ$ from 0.650 to 0.902 and agrees with the held-out additive-model oracle in 100\% of per-language decisions versus 68.5\% raw; these are consistency diagnostics, not human-grounded correctness measures. On a separately collected M-RewardBench panel (7 languages, 1{,}500 items per language, 10{,}500 language-item instances, 5 evaluators), panel agreement with the public human gold preferences rises from 68.7\% to 76.6\% (gain 7.9 percentage points, 95\% CI $[6.0, 9.9]$), our strongest external evidence of downstream usefulness. The estimator is the standard two-way ANOVA interaction-recovery operation under sum-to-zero contrasts; our contribution is its application as a label-free post-hoc calibrator for multilingual LLM judges, an explicit finite-sample concentration bound, and an unbiasedness result that holds even under task-language misspecification.

↑