Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
可扩展评估的极限:LLM作为裁判无法超越两倍数据
机构 * Max Planck Institute for Intelligent Systems, Tübingen(马克斯·普朗克智能系统研究所,图宾根) ; Tübingen AI Center(图宾根人工智能中心) ; ETH Zürich(苏黎世联邦理工学院)
AI总结 本文研究了使用LLM作为裁判进行模型评估的局限性,发现当裁判准确性不足时,去偏方法无法显著减少所需的真实标签数量。
Comments ICLR 2025; 27 pages, 8 figures