以人类为锚定的大语言模型裁判对新模型排名推断
Human-Anchored Inference for Ranking New Models with Large Language Model Judges
另 1 家 · 查看机构详情
- University of Science and Technology of China(中国科学技术大学)
- University of Minnesota(明尼苏达大学)
- National University of Singapore(新加坡国立大学)
- Peking University Health Science Center, Peking University(北京大学医学部,北京大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对新模型缺乏人类比较而仅有LLM裁判比较的排名问题,提出ANCHOR方法,利用历史数据学习裁判偏差并推断人类参考分数,实现高效准确排名。
中文摘要 AI 辅助
人类成对比较为评估大型语言模型(LLMs)提供了参考,但为每个新发布的模型收集足够的判断既昂贵又耗时。LLM裁判提供了一种可扩展的替代方案,尽管它们的比较可能系统地偏离人类偏好,且在不同裁判之间存在差异。我们研究一个已获得LLM裁判比较但缺乏人类比较的新模型的排名问题。我们提出ANCHOR(基于正交Riesz校正的锚定比较用于人类参考推断),该方法利用历史人类和LLM比较来学习裁判对人类分数差异的特定敏感性和特征相关的裁判偏差。这些估计随后用于从新模型的裁判比较中推断其人类参考分数。该框架允许特征分布从历史比较到新模型比较发生变化。为了进行推断,我们通过联合Riesz校正构建了一个Neyman正交估计器,该校正消除了估计历史人类分数、裁判敏感性和偏差函数的一阶效应。我们确立了识别性、收敛速率和具有一致可估计方差的新近正态性,并表明ANCHOR达到了半参数效率界。模拟实验展示了在分数估计和排名准确性方面的提升。在Chatbot Arena上,ANCHOR在竞争方法中实现了最低的分数均方根误差(RMSE)和插入平均绝对误差(MAE),且平均分数区间更窄。
英文摘要
Human pairwise comparisons provide a reference for evaluating large language models (LLMs), but collecting sufficient judgments for each new release is costly and time-consuming. LLM judges offer a scalable alternative, although their comparisons may differ systematically from human preferences and across judges. We study the ranking of a new model that has received LLM-judge comparisons but no human comparisons. We propose ANCHOR (ANchored Comparisons for Human-reference inference with Orthogonal Riesz correction), which uses historical human and LLM comparisons to learn judge-specific sensitivities to human score differences and feature-dependent judge biases. These estimates are then used to infer the new model's human-reference score from its judge comparisons. The framework allows the feature distribution to change between historical and new-model comparisons. For inference, we construct a Neyman-orthogonal estimator through a joint Riesz correction that removes the first-order effects of estimating the historical human scores, judge sensitivities, and bias functions. We establish identification, convergence rates, and asymptotic normality with consistently estimable variance, and show that ANCHOR attains the semiparametric efficiency bound. Simulations demonstrate gains in score estimation and ranking accuracy. On Chatbot Arena, ANCHOR achieves the lowest score RMSE and insertion MAE among competing methods, with narrower score intervals on average.