arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11947cs.CLcs.AI

无标签策略下准确率与顺序敏感性存在分歧

Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

Karl Hanna, Chen Feng

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对大语言模型的多项选择题评估偏差,测试两种无标签缓解策略,发现其无法可靠提升准确率,循环排列可提升准确率,且去偏差度量未达预期效果。

中文摘要 AI 辅助

多项选择题(MCQ)基准被广泛用于评估大语言模型,但MCQ得分将知识与对选项顺序的敏感性混为一谈,这使得它们无法可靠衡量模型的知识水平。本文中,我们测试让模型在给出答案时不查看选项标签,是否能消除位置影响并进而提升性能。我们评估了两种不同的缓解偏差策略:第一种采用生成后匹配方法,第二种是对选项单独打分,该方法在构造上无位置偏差。两种策略均无法可靠提升准确率。完整分解结果显示,瓶颈在于隐瞒选项,而非匹配步骤。唯一能始终匹配基线的配置是向模型展示所有选项并搭配大语言模型(LLM)匹配器。然而,完全消除位置影响仍无法可靠带来准确率提升,而循环排列常能提升准确率。对于两阶段提示,召回不平衡的聚合度量和针对每个问题的顺序敏感性直接度量均未显示出可靠的去偏差效果。

英文摘要

Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge. In this paper, we test whether preventing a model from seeing option labels while committing to an answer removes positional influence and, in turn, improves performance. We evaluate two different strategies for mitigating bias. The first uses a generation-then-matching approach, and the second scores options in isolation, which is positionally unbiased by construction. Neither reliably improves accuracy. A complete decomposition shows that the bottleneck is withholding options, not the matching step. The only configuration that consistently matches the baseline is the one that shows the model all options paired with an LLM matcher. However, eliminating positional influence entirely still does not reliably yield accuracy gains, while cyclic permutation often improves them. For two-stage prompting, an aggregate measure of recall imbalance and a direct per-question measure of order sensitivity both fail to show reliable debiasing.

发表机构

  • Queen’s University Belfast(贝尔法斯特女王大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑