利用按键级编辑特征预测CS1编程中的学习困难学生
Predicting Struggling Students in CS1 Programming Using Keystroke-Level Editing Features
浏览论文内容
中文总结 AI 辅助
该研究利用CodeBench平台的按键级编辑特征,在练习早期预测CS1编程中完全停滞的学生,发现组合特征的预测效果最优,初期信号最强。
中文摘要 AI 辅助
本文探究了在CS1编程练习期间利用按键级日志早期检测学习困难学生的可行性。部分学生在练习结束前无法得出正确解决方案,而当成绩或最终结果显现这一情况时,教师提供及时支持以帮助他们补救的机会可能已经错失。我们使用来自CodeBench平台的数据,该平台以按键级记录实时代码编辑事件,同时还包含执行日志和提交日志。我们定义了两个结果组:突破(BT)学生,其所有前期提交均得0%,最终提交获得满分;完全停滞(FS)学生,其所有提交均得0%且未得出正确解决方案。为检验该可行性,我们聚焦两个研究问题:(RQ1)相较于仅使用执行日志特征,添加按键级编辑特征是否能提升对FS学生的预测效果;(RQ2)在哪个阶段能最准确地预测BT和FS学生。对2019-1学期CodeBench数据集(包含507名学生)开展实验,比较三种特征配置:仅基于执行的特征(ExecOnly)、仅基于CodeMirror的特征(CMOnly)以及二者的组合(Combined)。我们在每次练习期间基于提交的连续阶段评估预测效果。在最早阶段,CMOnly的表现优于ExecOnly(AUROC为0.654 vs. 0.575),Combined较ExecOnly进一步提升了0.098(AUROC为0.674)。在所有配置中,最早阶段产生的预测信号最强。这些发现表明,练习初期存在的行为信号包含学生最终能否解决问题的有用线索,且按键级编辑日志相较于仅执行日志,为早期优先级排序提供了额外价值。
英文摘要
This paper investigates the feasibility of early detection of struggling students during CS1 programming exercises using keystroke-level logs. Some students fail to reach a correct solution before the exercise ends, and by the time this becomes apparent from grades or final outcomes, the opportunity for timely instructor support aimed at helping them recover may have passed. We use data from the CodeBench platform, which records real-time code editing events at the keystroke level, alongside execution and submission logs. We define two outcome groups: Breakthrough (BT) students, whose prior submissions all receive 0% and whose final submission achieves full credit, and Fully Stuck (FS) students, whose submissions all receive 0% without reaching a correct solution. To examine this feasibility, we focus on two questions: (RQ1) whether adding keystroke-level editing features improves the prediction of FS students over execution-log features alone, and (RQ2) at which stage BT and FS students can be predicted most accurately. Experiments on the 2019-1 semester of the CodeBench dataset, comprising 507 students, compare three feature configurations: execution-based features (ExecOnly), CodeMirror-based features (CMOnly), and their combination (Combined). We evaluate prediction across successive submission-based stages during each exercise. In the earliest stage, CMOnly outperforms ExecOnly (AUROC 0.654 vs. 0.575), and Combined further improves over ExecOnly by +0.098 (AUROC 0.674). Across all configurations, the earliest stage yielded the strongest predictive signal. These findings indicate that behavioral signals present at the very start of an exercise contain useful clues about whether a student will ultimately solve the problem, and that keystroke-level editing logs provide additional value for early prioritization beyond execution logs alone.