发表机构
George Mason University; Colgate University(乔治梅森大学; 科尔盖特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出将单选题质量保障视为一系列独立验证的决策,通过综述自动检测与修订方法,强调标准化的报告和结果导向的评估。
AI 中文摘要
生成单选题的规模日益扩大,但确定其评估质量仍然困难。我们呈现了一篇聚焦的叙述性综述,涵盖自动题目缺陷检测、修订、心理测量学筛选以及自然语言处理基准审计。通过数据库检索、引文检索和提名来源,我们获取了十四份研究报告并进行了全文审阅。我们区分了表面检查与内容敏感判断,并将一个包含19项标准的评分细则映射到检测方法和报告的证据上。高标签级准确率常常与弱阳性案例检测并存,而评分细则定义和参考标准各不相同。修订证据混杂,先前工作中报告的关联并未确立修复的效果。我们提议将质量保障评估为一系列独立验证的决策序列,并采用针对具体标准的报告、校准的人工审查以及基于结果的修订评估。
英文摘要
Generating multiple-choice questions is increasingly scalable, but establishing their assessment quality remains difficult. We present a focused narrative review of automated item-writing flaw detection, revision, psychometric screening, and NLP benchmark auditing. Database searches, citation retrieval, and nominated sources yield fourteen research reports reviewed in full text. We distinguish surface checks from content-sensitive judgments and map a 19-criterion rubric to detection methods and reported evidence. High label-level accuracy often coexists with weak positive case detection, while rubric definitions and reference standards vary. Revision evidence is mixed, and the associations reported in prior work do not establish the effects of repair. We propose evaluating quality assurance as a sequence of independently validated decisions, with criterion-specific reporting, calibrated human review, and outcome-based assessment of revisions.
Comments8 pages, 2 tables, Full paper accepted to the AIME Conference 2026