arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

土耳其 MMLU Pro:可追踪选项增强及其在土耳其语多项选择评估中的有效性边界

Turkish MMLU Pro: Traceable Option Augmentation and Its Validity Limits in Turkish Multiple-Choice Evaluation

M. Ali Bayram

arXiv 2609.15467首次发表:更新:

发表机构

Yıldız Technical University(耶尔德兹技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过构建土耳其 MMLU Pro 数据集,系统分析添加选项对多项选择评估的影响,揭示分数下降源于选项干扰而非知识测量改进,并明确该构建的有效性边界。

AI 中文摘要

添加答案选项可以降低多项选择分数,但并未提高评估的有效性。土耳其 MMLU Pro 使用横跨 58 个章节的 12,000 道土耳其语来源问题来检验这一区别。每个问题保留其题干、五个原始选项和来源键,并接收从同一章节其他问题中复制来的五个选项。句子嵌入检索提出候选选项;语言模型选择现有标识符。确定性验证重构所有 60,000 个添加项。一个包含 25 个模型的校准实验揭示了评分和生成预算效应。五次评估产生 34.8% 至 81.4% 的来源键准确率。在 981 个共享问题上,一个 API 服务的模型从五个选项时的 93.7% 下降到十个选项时的 83.1%;115 个丢失的正确回答中有 102 个选择了借用选项。在启发式标记的否定词干上,下降幅度为 24.4 个百分点,在其他地方为 5.9 个百分点。一项已完成的人工检查审计覆盖 200 个抽样问题,且审查者工具使用未记录,在两个记录集中分别产生 47 个和 31 个多答案判断,其中后者有 25 个未解决。这些记录支持对模糊性的担忧,而它们的依赖性和不完整的审查者方法文档限制了验证。由于顺序和标签也会改变,配对比较衡量的是实际实施的增强。贡献在于可追踪的构建及其有效性边界的分析,而非证明较低的十选项分数能更好地衡量知识。

英文摘要

Adding answer options can lower multiple-choice scores without improving assessment validity. Turkish MMLU Pro examines this distinction using 12,000 Turkish-source questions across 58 sections. Each question retains its stem, five original options and source key, and receives five options copied from other questions in the same section. Sentence-embedding retrieval proposes candidates; a language model selects existing identifiers. Deterministic verification reconstructs all 60,000 additions. A 25-model calibration exposes scoring and generation-budget effects. Five evaluations produce source-key accuracies of 34.8%-81.4%. On 981 shared questions, one API-served model falls from 93.7% with five choices to 83.1% with ten; 102 of 115 lost correct responses select borrowed options. The decrease is 24.4 percentage points on heuristically flagged negative stems and 5.9 points elsewhere. A completed human-checked audit of 200 sampled questions, with undocumented reviewer tool use, yields 47 and 31 multiple-answer judgments across the two record sets, 25 of the latter unresolved. These records support concern about ambiguity, while their dependence and incomplete reviewer-method documentation limit validation. Because order and labels also change, the paired comparison measures augmentation as implemented. The contribution is a traceable construction and an analysis of its validity limits, not evidence that lower ten-choice scores measure knowledge better.

CommentsPreprint of a manuscript submitted to Transactions on Machine Learning Research (TMLR)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑