JudgeArena:用于可复现LLM-评判者评估的统一框架
JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation
浏览论文内容
中文总结 AI 辅助
JudgeArena是统一主流LLM评判者基准的开源框架,支持可替换评判者,配备调优开源评判者配置,可高精度模拟LMArena Elo分数,减少对闭源模型的依赖,提升评估的透明度与可复现性。
中文摘要 AI 辅助
LLM作为评判者的评估已成为对语言模型进行排名的主流范式,但该生态系统仍处于碎片化状态:大多数基准测试附带各自的代码库,硬编码特定的闭源模型评判者,且仅支持单一评估协议。这种碎片化使得难以研究基准测试、评判者模型、提示词、推理后端等设计选择如何影响我们对模型质量得出的结论。我们推出JudgeArena,这是一个开源框架,它将主要的LLM评判者基准测试(AlpacaEval、Arena-Hard、MT-Bench和m-Arena-Hard)统一到单一接口下,支持可替换的评判者,并记录全面的元数据,以提高报告的透明度和可复现性。它支持对评判者选择进行系统研究,因为任何可通过vLLM、this http URL或OpenRouter访问的模型都可同时作为候选模型和评判者。此外,JudgeArena配备了针对开源模型的调优评判者配置,这些配置在英语和多语言设置的人类偏好数据集上得到验证,其性能与闭源模型评判者相当或更优,减少了对不透明闭源模型的依赖。最后,通过结合现有的人类标注和目标模型的LLM评判者评估,JudgeArena可以高精度模拟LMArena Elo分数,为大规模人类标注活动提供了一种实用、开源且低成本的替代方案。
英文摘要
LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their own code base, hardcode a specific closed-model judge, and support a single evaluation protocol. This fragmentation makes it difficult to study how design choices--the benchmark, the judge model, the prompt, the inference backend--affect the conclusions we draw about model quality. We introduce JudgeArena, an open-source framework that unifies major LLM-judge benchmarks (AlpacaEval, Arena-Hard, MT-Bench, and m-Arena-Hard) under a single interface with swappable judges and comprehensive metadata logging for increased transparency in reporting and reproducibility. It enables systematic studies of judge choices, as any model accessible via vLLM, llama.cpp, or OpenRouter can serve as both candidate and judge. Furthermore, JudgeArena ships with tuned judge configurations for open models that match or outperform closed-model judges, validated on human preference datasets in both English and multilingual settings, reducing the reliance on opaque closed models. Finally, by combining existing human annotations with LLM-judge evaluations of a target model, JudgeArena can simulate LMArena Elo scores with high accuracy offering a practical, open, and low-cost alternative to large-scale human annotation campaigns.
发表机构
- University of Freiburg(弗莱堡大学)
- ELLIS Institute Tübingen(图宾根ELLIS研究所)
- Microsoft AI(微软人工智能部门)
- Cohere Labs(Cohere实验室)
- Prior Labs(Prior实验室)
机构由 AI 辅助整理,请以论文原文为准。