arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型系统性偏向热门选项:多项选择题中的证据与缓解方法

Large Language Models Systematically Favor Popular Options: Evidence and Mitigation Across MCQs

Abdelrahman Abdallah, Mohammed Ali, Bhawna Piryani, Mahmoud Abdalla, Adam Jatowt

arXiv 2608.29257首次发表:更新:

发表机构

University of Innsbruck; Chungbuk National University(因斯布鲁克大学; 忠北国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLMs在MCQ评估中存在的流行度偏差问题,构建了PopMCQ基准,提出PopDebias修正方法,在22个开源LLMs上实现最高54.1个百分点的准确率提升。

AI 中文摘要

多项选择题(MCQs)是评估大语言模型(LLMs)的标准格式,但答案选项的流行度会干扰评估。现代LLMs系统性地偏好热门但错误的选项,而非较不热门的正确选项,我们将这种漏洞称为“流行度偏差”。这种模式与置信度校准错误一致:即使热门选项的准确率下降,模型的置信度仍保持较高水平。为系统分离该现象,我们引入了PopMCQ,这是一个包含六种受控策略的基准,这些策略在保持正确答案固定的同时改变选项的流行度。在我们最具对抗性的设置中,所有干扰项都比正确选项更受欢迎,模型选择热门错误答案的概率为66%。为缓解这种偏差,我们提出了PopDebias,这是一种轻量级推理时修正方法,用于从模型预测中估计并去除流行度先验。它无需微调,测试时无标签(仅使用小型校准拆分进行参数拟合),且计算开销可忽略不计。对22个开源LLMs(参数规模从0.5B到32B)的实验显示出一致的改进,在强流行度压力下,准确率提升高达54.1个百分点。代码和数据可在该https链接获取。

英文摘要

Multiple-choice questions (MCQs) are a standard format for evaluating large language models (LLMs), yet the popularity of answer options can confound evaluation. Modern LLMs systematically prefer popular but incorrect options over less popular correct ones, a vulnerability we call \textbf{popularity bias}. This pattern aligns with confidence miscalibration: model confidence remains high even as accuracy collapses for popular options. To systematically isolate this phenomenon, we introduce \textbf{PopMCQ}, a benchmark with six controlled strategies that vary option popularity while keeping the correct answer fixed. In our most adversarial setting, where all distractors are more popular than the correct option, models choose popular but wrong answers 66\% of the time. To mitigate this bias, we propose \textbf{PopDebias}, a lightweight inference-time correction that estimates and removes a popularity prior from model predictions. It requires no fine-tuning, is label-free at test time (using only a small calibration split for parameter fitting), and adds negligible computational cost. Experiments on 22 open-source LLMs (0.5B to 32B parameters) show consistent improvements, with accuracy gains up to 54.1 percentage points under strong popularity pressure. The code and data are available https://github.com/DataScienceUIBK/PopMCQ

CommentsAccepted at MAIN EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑