AI 中文总结
OSCAR提出顺序感知评分与校准框架,通过建模位置等呈现效应改进AI排名,实验显示显著提升预测精度与覆盖率。
AI 中文摘要
在聚合成对LLM评估时,评判者特异性敏感性是有用的,但其解释取决于排名模型包含哪些系统性呈现效应。我们引入了OSCAR,一个用于AI排名评分与校准的顺序感知框架,并将位置作为此类效应之一进行研究。在18位评估者发布的评判中,全响应A减B分数差异范围从-63.11到98.31个百分点。在已发布表格中匹配问题文本、响应文本、候选身份和评判者,在给定已发布文本映射的条件下,总体差异为24.22个百分点(95%区间[22.90,25.54])。一项受控计算隔离了潜在后果:当真实敏感性固定为1时,省略位置截距4会将总体最优斜率降至0.0771。我们扩展了基于敏感性的排名,加入评判者特定的位置、长度和族项,刻画了局部省略引起的位移和识别失败,并将提示簇不确定性传播到调整后的比较中。在四个已发布数据集上,位置提供了最大的独立预测改进。重新拟合的bootstrap比较显示,完整模型相对于仅位置调整有更具选择性的增益。在依赖二元模拟中,同时调整均值和协方差可获得94.4%至95.2%的覆盖率;仅修正其中任一单独项是不够的。在N=10,000时,OSCAR将平均中性目标RMSE从仅敏感性模型的0.1158降至0.0237。
英文摘要
Judge-specific sensitivity is useful for aggregating pairwise LLM evaluations, but its interpretation depends on which systematic presentation effects the ranking model includes. We introduce OSCAR, an order-aware framework for scoring and calibrating AI rankings, and study position as one such effect. In released judgments from 18 evaluators, the all-response A-minus-B score difference ranges from $-63.11$ to $98.31$ percentage points. Matching question text, response texts, candidate identities, and judge within the released table gives an overall difference of $24.22$ points (95% interval $[22.90,25.54]$), conditional on the released text mapping. A controlled calculation isolates the potential consequence: with true sensitivity fixed at one, omitting a position intercept of four reduces the population-optimal slope to $0.0771$. We extend sensitivity-based ranking with judge-specific position, length, and family terms, characterize local omission-induced displacement and an identification failure, and propagate prompt-cluster uncertainty to adjusted comparisons. Across four released datasets, position provides the largest stand-alone predictive improvement. Refitting bootstrap comparisons show more selective gains from the full model over position-only adjustment. In dependent binary simulations, adjusting both the mean and covariance yields 94.4--95.2% coverage; correcting either alone is insufficient. At $N=10{,}000$, OSCAR reduces mean neutral-target RMSE from $0.1158$ under the sensitivity-only model to $0.0237$.