arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型课堂投票语料库中测量质量与答案偏见的审计

An Audit of Measurement Quality and Answer Bias in a Large Classroom-Poll Corpus

Rohit Sharma, Pavani Ayinampudi, Aditya B. M. V., Jinal Gupta, Prakash Hegade, Sakshi Sharma, Meenakshi V, SRS Iyengar

arXiv 2609.38959首次发表:更新:

发表机构

Indian Institute of Technology Ropar(印度理工学院罗帕尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究审计了包含604个真实课堂投票的语料库,发现其测量精度中等且约四分之一题目区分度低,并揭示了学生普遍偏向回答“真”的答案偏见,该偏见影响题目难度,且可通过现有投票系统信号在题目生成时进行检测。

AI 中文摘要

实时课堂投票被广泛使用,且越来越多地通过自动化辅助生成,然而这些问题本身很少作为测量工具被评估。我们对一个大型真实课堂投票语料库进行了审计,该语料库包含47个课堂环节中的604个问题,被2,807名学习者回答了340,668次,将其作为一个测量工具。对于539个其正确答案能够从讲座记录中确定并验证的问题,我们使用项目反应理论将每个问题和每个学生置于同一量表上,并分析了真/假题目的答案结构。研究发现两个结果。第一,这些投票形成了一个连贯但容易的量表,精度中等(边际信度约为0.60),在该量表上,大约四分之一的问题几乎无法区分较强和较弱的学生。第二,学生表现出一种稳健的倾向于回答“真”的倾向,这在个体层面存在(77%的学生偏向于回答“真”),这与题目倾向于以“假”为正确答案的较温和倾向相遇;因此,答案方向预测了难度,以“假”为正确答案的题目难度大约高出十三个百分点,并且该效应在控制题目内容和选择性作答后依然存在。这两个发现都依赖于投票系统已经记录的信号,因此同样的检查可以在题目生成时、在它们到达学生之前运行。

英文摘要

Real-time classroom polls are widely used and increasingly generated with automated assistance, yet the questions themselves are rarely evaluated as measurements. We audit a large corpus of authentic classroom polls, 604 items across 47 sessions answered 340,668 times by 2,807 learners, as a measurement instrument. For the 539 items whose correct answer could be established and verified from the lecture transcript, we place every item and every student on a common scale using item response theory and analyse the answer structure of the True/False items. Two findings emerge. First, the polls form a coherent but easy scale of moderate precision (marginal reliability about 0.60), on which roughly a quarter of items barely separate stronger from weaker students. Second, students show a robust tendency to answer True, present at the individual level (77% of students lean True), which meets a milder tendency for items to be keyed False; as a result answer direction predicts difficulty, False-keyed items being about thirteen points harder, and the effect survives controls for item content and for selective answering. Both findings rest on signals a polling system already records, so the same checks can be run as items are generated, before they reach students.

Comments14 pages, 6 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑