基于多尝试编程轨迹的感知修订的提交成功预测
Revision-Aware Success Prediction from Multi-Attempt Programming Trajectories
浏览论文内容
中文总结 AI 辅助
本研究针对编程轨迹的三项预测任务,系统比较ML、DL及PTM模型的表现,发现ML模型整体最优且稳定,为编程教育的提交感知分析提供了建模指导。
中文摘要 AI 辅助
编程结果预测在数据驱动的编程教育中占据核心地位,可支持学习者建模、及时干预和自适应辅助。然而,由于编程轨迹中存在异构错误状态、短期修订以及未来时间范围的可用性不均,预测提交成功颇具难度。本研究在统一框架下考察三项预测任务:当前尝试是否被接受(任务1)、下一次尝试是否被接受(任务2)、在三次尝试的恢复窗口内是否达到接受状态(任务3)。每个任务均采用线性支持向量机(LinearSVM)、XGBoost、双向门控循环单元(BiGRU)、双向长短期记忆网络(BiLSTM)、GraphCodeBERT和CodeT5+等机器学习(ML)、深度学习(DL)及基于Transformer的预训练模型(PTM),在仅当前输入、成对输入和多步输入三种场景下进行评估。结果显示存在一致规律:仅当前输入场景最可靠,而成对和多步历史输入未带来一致增益;整体而言,ML模型是最强且最稳定的,尤其在任务1和任务3中表现突出,且任务2在所有模型类别中均最难;DL和PTM在任务3中表现良好,但更具任务依赖性。在任务3的仅当前输入设置下,XGBoost的平均精度(AP)/精度-召回曲线下面积(PR-AUC)达99.09%,马修斯相关系数(MCC)为0.6325,而GraphCodeBERT和CodeT5+的F1分数分别为80.00%和73.68%。敏感性分析证实,在更严格的未来时间范围控制下,ML模型对任务3的结论最具稳健性。在所有设置中,ML模型对编程成功预测仍高度有效,而复杂模型在特定场景下具有价值。本研究对不同预测框架进行了系统比较,为编程教育中的提交感知分析提供了稳健的建模指导,其中近期成功预测可用于在线评测平台的及时干预和自适应编程支持系统。
英文摘要
Programming outcome prediction plays a central role in data-driven programming education, supporting learner modeling, timely intervention, and adaptive assistance. Yet predicting submission success is difficult due to heterogeneous error states, short-term revisions, and uneven future-horizon availability in programming trajectories. This study examines three prediction tasks under a unified formulation: whether the current attempt is accepted (Task~1), whether the next attempt is accepted (Task~2), and whether acceptance is reached within a three-attempt recovery window (Task~3). Each task is evaluated across current-only, pairwise, and multi-step input regimes using ML, DL, and transformer-based pretrained models (PTM), represented by LinearSVM, XGBoost, BiGRU, BiLSTM, GraphCodeBERT, and CodeT5+. Results show a consistent pattern: the current-only regime is the most reliable, while pairwise and multi-step history provide no consistent gain. ML models are the strongest and most stable overall, particularly in Tasks~1 and~3, and Task~2 is the hardest across all model families. DL and PTMs perform well on Task~3 but are more task-dependent. In the Task~3 current-only setting, XGBoost achieves AP/PR-AUC of 99.09% and MCC of 0.6325, while GraphCodeBERT and CodeT5+ reach F1 scores of 80.00% and 73.68%, respectively. A sensitivity analysis confirms that Task~3 conclusions hold most robustly for ML models under stricter future-horizon control. Across all settings, ML models remain highly effective for programming success prediction, while complex models offer value in specific settings. This work provides a systematic comparison across predictive formulations and offers robust modeling guidance for submission-aware analytics in programming education, where near-future success prediction can inform timely intervention in online judge platforms and adaptive programming support systems.