arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JudgeMoE:面向大语言模型裁判的分布聚合方法

JudgeMoE: Distributional Aggregation for LLM-as-a-Judge

Yiqi Liu, Joseph James, Yang Wang, Kun Zhao, Chenghao Xiao, Chenghua Lin

arXiv 2610.07109首次发表:更新:

发表机构

The University of Manchester; The University of Sheffield; University of Pittsburgh; Shanghai University of Finance and Economics(曼彻斯特大学; 谢菲尔德大学; 匹兹堡大学; 上海财经大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

JudgeMoE通过为缓存的大语言模型裁判分数分布分配样本特定权重并融合,解决了标量压缩丢失信息的问题,在16个单元上平均提升0.0393,优于最强单裁判。

AI 中文摘要

当大语言模型裁判对输出进行评分时,其分数分布保留了在标量压缩后丢失的不确定性和分歧信息。我们提出了JudgeMoE,一种轻量级聚合器,它为缓存的裁判分数分布分配特定于样本的权重,并在计算最终分数之前进行融合。一项协议研究表明,分数范围的选择在裁判-数据集设置中不稳定,且软评分通常优于硬解码。在原始的10单元基准上,JudgeMoE将平均Spearman相关系数相对于均匀对数池化提升了+0.079。将相同配置应用于另外六个单元,在16个单元上相对于最强的单个局部裁判获得了+0.0393的平均增益,其中12/16个单元存在正差异,单侧Wilcoxon符号秩检验p=0.0091。基于验证的分析进一步表明,优选的聚合方法取决于任务和裁判池。

英文摘要

When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression. We introduce JudgeMoE, a lightweight aggregator that assigns example-specific weights to cached judge score distributions and fuses them before computing a final score. A protocol study shows that score-range choice is unstable across judge--dataset settings and that soft scoring usually outperforms hard decoding. On the original 10-cell benchmark, JudgeMoE improves mean Spearman over uniform log pooling by $+0.079$. Applying the same configuration to six additional cells yields a $+0.0393$ mean gain over the strongest local single judge across 16 cells, with positive differences in 12/16 cells and a one-sided Wilcoxon signed-rank $p=0.0091$. Validation-based analyses further show that the preferred aggregation method depends on the task and judge pool.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑