使用AI评审进行可信方法比较:顺序、批次和聚合效应下的估计与设计
Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects
浏览论文内容
中文总结 AI 辅助
本研究提出用马尔可夫广义线性混合模型近似LLM评审机制,分析排行榜排名与组间比较的统计性质,并验证了模型在多种设置下的有效性。
中文摘要 AI 辅助
大型语言模型(LLMs)越来越多地被用作自动化AI评估的评审。常见做法是随机化提示顺序并平均所得分数,但其统计有效性尚不明确。我们证明LLM评估机制可以用一类马尔可夫广义线性混合模型(GLMMs)来近似,并得到三个主要商业LLM的样本外预测支持。使用一阶马尔可夫GLMM,我们研究排行榜排名和组间比较。对于排行榜排名,随机化-平均选择在温和的分离条件下是一致的,而当项目质量接近时,Williams方设计可以提高效率。对于组间比较,由于响应模型的非线性,朴素平均可能对组级质量差异得出不一致的结论。实证结果进一步支持所提出的基于模型的推断在一阶理论之外的有效性,包括具有高阶序列记忆的设置。我们在一个AI评审比较两种图形模型估计方法的应用中展示了该方法。
英文摘要
Large language models (LLMs) are increasingly used as judges for automated AI evaluation. A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear. We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear mixed models (GLMMs), supported by out-of-sample predictions across three major commercial LLMs. Using a first-order Markov GLMM, we study leaderboard ranking and group comparison. For leaderboard ranking, randomize-and-average selection is consistent under a mild separation condition, and a Williams square design can improve efficiency when item qualities are close. For group comparison, naive averaging can yield inconsistent conclusions about differences in group-level quality because of the response model's nonlinearity. Empirical results further support the validity of the proposed model-based inference beyond the first-order theory, including settings with higher-order sequence memory. We illustrate the approach in an application where AI judges compare two graphical model estimation methods.
发表机构
- University of Minnesota(明尼苏达大学)
机构由 AI 辅助整理,请以论文原文为准。