发表机构
Botswana International University of Science and Technology (BIUST); Stellenbosch University; DSI/NRF Centre of Excellence in STI Policy, Stellenbosch University; University of South Africa(博茨瓦纳国际科学与技术大学; 斯坦陵布什大学; 斯坦陵布什大学科技与创新政策卓越中心; 南非大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种基于信息论的多项选择题干扰项分析方法,通过熵衡量错误答案的随机程度,以识别集中误解并辅助教学诊断。
AI 中文摘要
多项选择题因其高效性而成为教育评估中的常用工具,但其诊断能力本质上受到限制。传统的心理测量技术通过将模型拟合到预定数据模式来推断误解,这使得它们在分析缺乏此类先验趋势的小型、新颖或非标准化样本时效果不佳。为解决这一问题,我们提出了一种基于信息论的新干扰项分析方法,该方法将错误答案中的随机性视为学生答案模式中可测量的信号,而非噪声。通过同时考察干扰项选择的熵和整体准确率,我们的方法提供了一个与样本大小无关的框架来对表现进行分类。将低分项目层面的响应数据与标准力学清单中独立记录的误解进行交叉对照,结果表明,对于被识别为最集中(随机程度最低)的项目,大多数错误响应都选择了与该项目命名且经访谈验证的误解相对应的确切选项。此外,在其他地方被独立标记为诊断可靠性不佳的项目,恰好是该测量方法识别为具有最高随机程度的项目。在教学上可操作的情况是班级错误答案中随机程度较低的情况,这表明存在一种共同的、可重新教学的误解,而高随机程度主要作为判断这种集中性的零假设情况。我们的分析旨在作为一种快速、单次施测的分诊步骤:一种实用且理论驱动的补充方法,与既有的干扰项分析方法相辅相成,帮助教师标记哪些项目值得进行更密切、更耗费资源的调查。
英文摘要
Multiple-choice questions are a staple of educational assessment due to their efficiency, but their diagnostic ability is inherently limited. Traditional psychometric techniques infer misconceptions by fitting models to predetermined data patterns, making them ineffective for analysing small, novel, or non-standardised samples where such prior trends are absent. To address this, we introduce a new distractor analysis method based on Information Theory, which treats the randomness in incorrect answers not as noise but as a measurable signal in student answer patterns. By examining the entropy of distractor choices alongside overall accuracy, our approach offers a sample-sized-independent framework for classifying performance. Cross-referencing low-scoring item-level response data against independently documented misconceptions on a standard mechanics inventory shows that, for items identified as most concentrated (lowest degree of randomness), a majority of incorrect responses select the exact response coded to that item's named, interview-validated misconception. Furthermore, items independently flagged elsewhere as diagnostically unreliable are exactly the items the measure identifies as having the highest degree of randomness. The pedagogically actionable cases are those with low degree of randomness in a class's incorrect answers, signalling a shared, re-teachable misconception, whereas a high degree of randomness chiefly serves as the null case against which that concentration is judged. Our analysis is intended as a fast, single-administration triage step: a practical, theory-driven complement to established distractor-analysis methods that helps instructors flag which items warrant closer, more resource-intensive investigation.