arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30372cs.CL

基于概率景观的MCQA基准测试审计

Auditing MCQA Benchmarks through Probability Landscapes

Minsoo Song, Chanjun Park

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出两部分概率框架,结合P_top1、H_norm、MPD及噪声注入方法,实现MCQA基准测试的基准级与条目级审计,其结果与专家注释一致,可作为轻量级审计工具。

中文摘要 AI 辅助

随着大型语言模型的快速发展,标准多项选择问答(MCQA)基准测试的性能已趋近饱和。尽管学界已通过开发难度不断提升的数据集做出回应,但验证问题质量与过滤有缺陷的条目仍是一项劳动密集型工作。为提供可扩展的诊断方法,我们提出了由两部分组成的概率框架,用于利用模型输出分布对MCQA基准测试进行审计:第一,针对基准级分析,我们使用最高预测概率(P_top1)和归一化残差熵(H_norm)来表征概率景观,并通过成对距离均值(MPD)进行全局总结;第二,针对条目级诊断,我们引入噪声注入以减少有意义的干扰项竞争,从而能够标记候选条目以进行针对性人工审查,并对剩余失败模式进行分类。在四个MCQA基准测试上,我们的景观分析揭示了模型置信度和剩余选项竞争的基准级差异;与此同时,我们的噪声注入方法标记了潜在可处理的条目级问题,显示出与MMLU-Redux的专家错误注释的一致性。这些结果表明,我们基于概率的框架提供了一种轻量级审计工具,用于比较宏观层面的基准测试结构,并为针对性人工审查确定单个条目的优先级。

英文摘要

As Large Language Models rapidly advance, performance on standard multiple-choice question answering (MCQA) benchmarks is reaching saturation. While the community has responded by developing increasingly difficult datasets, validating question quality and filtering flawed items remains a labor-intensive process. To provide a scalable diagnostic approach, we propose a two-component probabilistic framework for auditing MCQA benchmarks using model output distributions. First, for benchmark-level analysis, we characterize the probability landscape using the top prediction probability ($P_{top1}$) and normalized residual entropy ($H_{norm}$), summarized globally by Mean Pairwise Distance (MPD). Second, for item-level diagnostics, we introduce noise injection to reduce meaningful distractor competition, enabling us to flag candidate items for targeted human review and categorize residual failure patterns. Across four MCQA benchmarks, our landscape analysis reveals benchmark-level differences in model confidence and residual option competition. Concurrently, our noise-injection method flags potentially actionable item-level issues, showing alignment with expert error annotations from MMLU-Redux. These results suggest that our probability-based framework provides a lightweight audit lens for comparing macro-level benchmark structure and prioritizing individual items for targeted human review.

发表机构

  • Soongsil University(崇实大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑