发表机构
Hanyang University(汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出惩罚框架下的无有效选项MCQA,通过移除正确选项并允许弃权,揭示了高准确率模型在无效选项下仍可能产生错误强制选择,暴露了标准评估未涵盖的可靠性缺陷。
AI 中文摘要
多项选择问答(MCQA)通常用于评估大语言模型,其假设所提供的选项之一是正确的,并通常使用答案选择准确率作为评估指标。然而,在实际部署中,用户或检索系统可能提供无效选项集,其中所列选项均不正确,而选择其中任何一个都可能产生下游成本。我们将此场景研究为惩罚框架下的无有效选项MCQA。利用MMLU-Pro的数学子集,我们移除标记的正确选项,允许模型选择剩余选项之一或输出弃权(不执行),并对无效的强制选择响应施加惩罚。我们进一步引入了正确条件分析,仅对模型最初回答正确的实例评估弃权行为。实验表明,高MCQA准确率并不能完全保证弃权的可靠性:即使在明确的无有效选项感知指令和基于惩罚的评分下,模型仍会对一部分最初正确的实例产生无效的强制选择响应。这些结果表明,惩罚框架下的无有效选项MCQA揭示了标准答案选择准确率未捕获的模型可靠性方面。
英文摘要
Multiple-choice question answering (MCQA) is commonly used to evaluate large language models under the assumption that one of the provided options is correct, typically using answer-selection accuracy. However, in real deployments, users or retrieval systems may provide invalid option sets in which none of the listed choices is correct, and selecting one of them may incur downstream cost. We study this setting as penalty-framed no-valid-option MCQA. Using the mathematics subset of MMLU-Pro, we remove the labeled correct option, allow models to either choose a remaining option or output ABSTAIN, and penalize invalid forced-choice responses. We further introduce correct-conditioned analysis, evaluating abstention only on instances that the model originally answered correctly. Experiments show that high MCQA accuracy does not fully guarantee abstention reliability: even under explicit no-valid-option-aware instructions and penalty-based scoring, models still produce invalid forced-choice responses for a subset of originally correct instances. These results show that penalty-framed no-valid-option MCQA reveals an aspect of model reliability not captured by standard answer-selection accuracy.
CommentsAccepted to AACL-IJCNLP 2026 Main Conference (Short Paper)