arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27063cs.CY

学生使用LLM与数据科学课程中AI生成题目难度的局限性

Student Use of LLMs and the Limits of AI-Generated Question Difficulty in Data Science Courses

Yuan An, Lei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过德雷塞尔大学数据科学课程的多源课堂研究,发现学生与LLM的互动随学期增加且存在过度依赖倾向,而LLM生成的MCQ难度评级与布鲁姆分类水平相关但无法预测实际题目难度,反映结构格式而非真实难度。

中文摘要 AI 辅助

本文介绍了一项多来源课堂研究,该研究在德雷塞尔大学的数据科学课程中于一个为期10周的学季进行。我们首先通过三项数据科学课程中的四次调查,研究了学生与大语言模型(LLM)互动的行为。其次,我们评估了由LLM为课堂内检索练习生成的多项选择题(MCQ)的结构效度。基于378道编写的问题(其中311道已部署,产生了7,888条学生回答),我们分析了LLM分配的难度评级是否与经验性题目难度相匹配。我们的研究表明,学生与LLM的互动在不同课程间存在差异,并随着学期推进而增加。尽管学生表达了高度满意度并报告节省了大量时间,但他们对深度学习益处的感知有所下降,且许多人注意到存在过度依赖的倾向。关于LLM生成的MCQ的难度评级,简单、中等和困难标签与其分配的布鲁姆分类学水平高度相关(Spearman ρ=0.90),这反映了共同生成的人工产物。然而,这两个指标均未能预测经验性题目难度(难度标签ρ=0.06;布鲁姆水平ρ=0.02)。这些评级反映的是问题的结构格式,而非其潜在难度。

英文摘要

This paper presents a multi-source classroom study conducted during a 10-week quarter in data science courses at Drexel University. We first investigate the behaviors of student engagement with large language models (LLMs) using four surveys across three data science courses. Second, we evaluate the construct validity of multiple-choice questions (MCQs) generated by an LLM for in-lecture retrieval practice. Based on 378 authored questions (311 deployed, producing 7{,}888 student responses), we analyze whether the difficulty ratings assigned by an LLM match empirical item difficulty. Our study shows that student engagement with LLMs varied across courses and increased over the term. Although students expressed high satisfaction and reported saving considerable time, their perception of deep learning benefits declined, and many noted a tendency toward over-reliance. Regarding the difficulty ratings of LLM-generated MCQs, the Easy, Medium, and Hard labels correlated closely with its assigned Bloom's Taxonomy levels (Spearman $ρ=0.90$), reflecting an artifact of co-generation. However, neither metric predicted empirical item difficulty (difficulty label $ρ=0.06$; Bloom level $ρ=0.02$). The ratings reflect the structural formatting of a question rather than its underlying difficulty.

↑