arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

长度偏差下的准确率和归一化准确率:分析、指南及贝叶斯替代方法

Accuracy and Normalized Accuracy under Length Bias: Analysis, Guidelines, and a Bayesian Alternative

Koen Oostermeijer

arXiv 2607.12767首次发表:更新:

AI 中文总结

多项选择基准测试存在长度偏差,常见归一化方法常过度校正。本文分析评分规则,引入贝叶斯准确率,它能消除线性长度效应,可直接替代似然评估,无需额外前向传递,在多场景下长度偏差更低。

AI 中文摘要

通过条件对数概率对候选完成项进行排序的多项选择基准测试存在长度偏差:由于对数概率是对令牌求和,在实际中较长答案相对于较短答案往往会受到惩罚。常见的缓解方法是按完成长度对分数进行归一化,但我们通过实证表明,这种启发式方法经常过度校正,反而导致对较长答案的偏向。我们首先分析这些评分规则,确定标准准确率和长度归一化准确率何时适用,以及它们的长度偏差如何取决于完成长度的分布。基于此分析,我们引入了贝叶斯准确率,这是一种评分规则,可在答案长度的显式先验下计算每个候选项的后验概率,从而消除线性长度效应。贝叶斯准确率可直接替代基于似然的多项选择评估,无需额外的前向传递,并且在跨基准测试和少样本设置中始终表现出比标准准确率和长度归一化准确率更低的实证长度偏差。

英文摘要

Multiple-choice benchmarks that rank candidate completions by conditional log-probability suffer from a length bias: because log-probabilities sum over tokens, longer answers tend to be penalized relative to shorter ones in practice. A common mitigation is to normalize scores by completion length, but we show empirically that this heuristic frequently over-corrects, introducing a bias toward longer answers instead. We first analyze these scoring rules, characterizing when standard and length-normalized accuracy are appropriate and how their length biases depend on the distribution of completion lengths. Motivated by this analysis, we introduce \emph{Bayesian accuracy}, a scoring rule that computes the posterior probability of each candidate under an explicit prior over answer length, thereby removing linear length effects. Bayesian accuracy is a drop-in replacement for likelihood-based multiple-choice evaluation, requires no additional forward passes, and consistently exhibits lower empirical length bias than both standard and length-normalized accuracy across benchmarks and few-shot settings.

CommentsAccepted at ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑