发表机构
The University of Manchester; The University of Sheffield; University of Pittsburgh; Shanghai University of Finance and Economics(曼彻斯特大学; 谢菲尔德大学; 匹兹堡大学; 上海财经大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
JudgeMoE通过为缓存的大语言模型裁判分数分布分配样本特定权重并融合,解决了标量压缩丢失信息的问题,在16个单元上平均提升0.0393,优于最强单裁判。
AI 中文摘要
当大语言模型裁判对输出进行评分时,其分数分布保留了在标量压缩后丢失的不确定性和分歧信息。我们提出了JudgeMoE,一种轻量级聚合器,它为缓存的裁判分数分布分配特定于样本的权重,并在计算最终分数之前进行融合。一项协议研究表明,分数范围的选择在裁判-数据集设置中不稳定,且软评分通常优于硬解码。在原始的10单元基准上,JudgeMoE将平均Spearman相关系数相对于均匀对数池化提升了+0.079。将相同配置应用于另外六个单元,在16个单元上相对于最强的单个局部裁判获得了+0.0393的平均增益,其中12/16个单元存在正差异,单侧Wilcoxon符号秩检验p=0.0091。基于验证的分析进一步表明,优选的聚合方法取决于任务和裁判池。
英文摘要
When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression. We introduce JudgeMoE, a lightweight aggregator that assigns example-specific weights to cached judge score distributions and fuses them before computing a final score. A protocol study shows that score-range choice is unstable across judge--dataset settings and that soft scoring usually outperforms hard decoding. On the original 10-cell benchmark, JudgeMoE improves mean Spearman over uniform log pooling by $+0.079$. Applying the same configuration to six additional cells yields a $+0.0393$ mean gain over the strongest local single judge across 16 cells, with positive differences in 12/16 cells and a one-sided Wilcoxon signed-rank $p=0.0091$. Validation-based analyses further show that the preferred aggregation method depends on the task and judge pool.