AI 中文总结
针对大语言模型传统评估基准不足,本文提出基于共识的评估框架,通过不同模型对候选响应排序,以模型间一致性衡量质量,经多模型跨领域研究揭示偏好模式,提供了可扩展的比较评估方法。
AI 中文摘要
传统的大语言模型基准主要依赖静态数据集和客观评分指标,在多个答案可接受时往往无法捕捉响应质量的差异。本文引入一种基于共识的评估框架,通过不同大语言模型对同一提示的匿名候选响应进行排序,将模型间的总体一致性视为盲目条件下感知响应质量的代理。通过对五个跨领域的最先进大语言模型进行控制研究,结果揭示了跨领域一致的偏好模式,此框架为比较评估提供了可扩展、模型驱动的方法。
英文摘要
Traditional benchmarks for LLMs primarily rely on static datasets and objective scoring metrics, which often fail to capture differences in response quality when multiple answers are acceptable. In such settings, correctness alone is insufficient to distinguish between responses that vary in clarity, completeness, and usefulness. This paper introduces a consensus-based evaluation framework that measures relative preference among model-generated responses rather than absolute correctness. Instead of evaluating outputs against a fixed ground truth, we assess how a panel of diverse LLMs ranks anonymized candidate responses to the same prompt. This approach treats aggregate inter-model agreement as a proxy for perceived response quality under blind conditions. We conduct a controlled study using five state-of-the-art LLMs across multiple domains, including programming, general knowledge, safety, logical reasoning, and mathematics. Each model generates responses and independently ranks peer outputs through a structured voting process. Scores are aggregated into a Relative Intelligence Index (RII), representing how frequently a model's responses are preferred by other models. Our findings reveal consistent preference patterns across domains, with certain models more frequently ranked highly by their peers. However, we emphasize that these results reflect inter-model preference alignment rather than objective correctness or human judgment. This framework provides a scalable, model-driven method for comparative evaluation, offering an alternative perspective on response quality in scenarios where multiple valid answers exist. While not directly aligned with human evaluation, prior work suggests that aggregated model preferences can partially correlate with human judgments, motivating this as a proxy signal.
Comments14 pages, 7 figures