AI 中文总结
本文评估开放对比决策模型CLM作为评判者的性能,发现其在公开基准上接近随机水平,但置信度校准和顺序稳定性有效,级联策略因置信度保留不足而受限。
AI 中文摘要
一个开放的对比决策模型在困难的公开基准上作为评判者时表现接近随机水平:Contrastive-LM/CLM-v0.1-8B的得分介于0.351(最佳四选一,随机水平为0.250)和0.593(两两比较,随机水平为0.500)之间,在RM-Bench和JudgeBench上与抛硬币的结果在统计上无法区分,并且对HaluEval的每个条目都回答同一个恒定标签,与平凡的始终选第一个的基线在0.581处持平。具有相同参数数量的评判者在所有地方得分都高得多:一个奖励模型达到0.764至0.976,一个生成式评判者达到0.611至0.778,并且在Benjamini-Hochberg校正后,与CLM的每个差距都是显著的。有两个特性确实有效。原始置信度过度自信,最高达+0.401,然而在保留的校准条目上拟合一个合并的温度参数,可将期望校准误差修复至最多0.062,并且修复后的置信度在六个基准中的三个上,对模型自身错误的排序优于随机水平。决策顺序翻转率为0.0002,而生成式评判者为0.2188,长度偏好偏移为-0.023,而生成式评判者为-0.217。然而,置信度门控级联在预注册的0.97保留阈值下,将0.923至1.000的条目升级到强评判者:关于一个接近随机水平的评判者的校准置信度几乎没有什么可以保留的。设计如下:五个公开偏好基准和一个带有真实标签的幻觉基准,在预注册冻结于看到任何测试条目之前的情况下进行评分,与生成式、奖励模型和普通基线进行比较,并发布每个条目的预测。
英文摘要
An open contrastive decision model is near chance as a judge on the hard public benchmarks: Contrastive-LM/CLM-v0.1-8B scores between 0.351 (best- of-four, chance 0.250) and 0.593 (pairwise, chance 0.500), is statistically indistinguishable from coin flipping on RM-Bench and JudgeBench, and answers every HaluEval item with one constant label, matching the trivial always-first baseline at 0.581. Judges with the same parameter count score far higher everywhere: a reward model reaches 0.764 to 0.976 and a generative judge 0.611 to 0.778, and every gap to CLM is significant after Benjamini-Hochberg correction. Two properties do work. Raw confidences are overconfident by up to +0.401, yet one pooled temperature fit on held-out calibration items repairs expected calibration error to at most 0.062, and the repaired confidence ranks the model's own errors above chance on three of six benchmarks. The decision order-flip rate is 0.0002 against 0.2188 for the generative judge, and the length-preference shift is -0.023 against -0.217. The confidence-gated cascade, however, escalates between 0.923 and 1.000 of items to the strong judge at the preregistered 0.97 retention bar: calibrated confidence about a near-chance judge has almost nothing to keep. The design: five public preference benchmarks and one hallucination benchmark with real labels, scored under a preregistration frozen before any test item was seen, against generative, reward-model, and trivial baselines, with per-item predictions released.