发表机构
King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RAIM通过交叉拟合堆叠回归聚合廉价开放权重评判者,以稳健替代前沿模型进行幻觉检测,在保留93%中位κ值的同时成本降至六十四分之一。
AI 中文摘要
对忠实性的自动评估日益依赖大型语言模型作为评判者,然而最可靠的评判者是专有的前沿模型,成本高昂且不适合高通量监控。我们研究是否可以将一组廉价的开放权重评判者(4-9B)聚合起来以替代前沿评判者,这种替代会牺牲什么,以及何时值得进行这种替代。我们提出了RAIM,一种对成员间相关误差具有稳健性的聚合方案,该方案将交叉拟合的堆叠逻辑回归与一个可接受性检验相结合,该检验从成员自身的输出中读取,用于判断聚合它们是否优于其最佳成员,并保持在接近前沿评判者的水平。我们使用来自八个忠实性基准测试中不同家族的十个评判者实例化RAIM。与Claude Sonnet相比,该面板保留了其中位Cohen's κ的93%,平均仅损失2.9个平衡准确率点;从配对差异来看,它在一个基准上明显改进,在三个基准上明显恶化(其中只有两个的差距不可忽略),其余四个未定。以前沿模型推理价格的六十四分之一,实际操作成本是对50-100个标注记录进行一次性域内校准。该面板在其主场上也与专门训练的检测器具有竞争力(与GPT-4o相差1.3个准确率点,与LLM-AggreFact领先者相差1.9个点),并在我们的接地集上比我们重新运行的最强检测器高出6个点。聚合是否值得取决于成员本身:在多个有能力的成员在不同项目上出错的地方,面板优于其最佳评判者并接近前沿;在一个成员占主导地位的地方,堆叠器恢复领先者,且只有在那里前沿模型才保持实质性领先。这两个条件都可以从校准集中读取,无需额外成本,因此,只要这种审计允许,廉价面板就可以替代前沿面板。
英文摘要
Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary frontier models, costly and ill-suited to high-throughput monitoring. We investigate whether a panel of cheap open-weight judges (4--9B) can be aggregated to stand in for a frontier one, what the substitution sacrifices, and when it is worth making. We propose RAIM, an aggregation scheme robust to the members' correlated errors, coupling a cross-fitted stacked logistic regression with an admissibility test that, read from the members' own outputs, identifies when aggregating them improves on their best member and stays within reach of the frontier judge. We instantiate RAIM with ten judges from disjoint families across eight faithfulness benchmarks. Against Claude Sonnet, the panel retains a median 93% of its Cohen's $κ$ and gives up only 2.9 points of balanced accuracy on average; read as paired differences, it clearly improves on one benchmark and clearly worsens on three (only two by a non-negligible margin), leaving four unresolved. At a sixty-fourth of the frontier's inference price, the operative expense is a one-time in-domain calibration on 50--100 labelled records. The panel is also competitive with purpose-trained detectors on their home benchmarks (within 1.3 accuracy points of GPT-4o and 1.9 of the LLM-AggreFact leader), and beats the strongest one we reran by 6 points on our grounded sets. Whether aggregation pays depends on the members themselves: where several capable members err on different items, the panel improves on its best judge and approaches the frontier; where one dominates, the stacker recovers the leader, and only there does the frontier remain materially ahead. Both conditions are read off the calibration set at no further cost, so a cheap panel can stand in for a frontier one wherever this audit admits it.
Comments49 pages, 23 tables, 10 figures. Code and data: https://github.com/eOnofri04/raim-analysis and https://github.com/eOnofri04/raim-verdicts