arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27586cs.CYcs.AI

解决生成式AI(GenAI)方案比评估更好吗?

Is Solving Better Than Evaluating GenAI Solutions?

Ethan Dickey, Marios Mertzanidis, Alexandros Psomas

首次发表
浏览论文内容

中文总结 AI 辅助

通过220名大三算法课学生的随机A/B交叉研究发现,评估GenAI方案与传统解题的整体成绩无显著差异,仅作业成绩局部更高,需刻意支架式教学才能获得学习收益。

中文摘要 AI 辅助

随着生成式AI(GenAI)工具生成计算作业解决方案的能力不断提升,计算教育界正在探索将方案评估、验证和批判与传统方案生成并重的教学方法。然而,这类以评估为中心的任务对学生学习的影响仍缺乏证据,尤其是在偏重理论的高年级课程中。我们在一门大三算法课程中开展了随机A/B交叉研究(N=220),对比评估GenAI生成的方案与传统解题的效果。在6次作业中,学生学习小组要么直接解决具有挑战性的算法问题,要么评估通常存在缺陷的GenAI生成方案,且学期中途两组角色互换。我们发现,两组在期中成绩、期末成绩、课程总绩点或与作业干预结构匹配的考试题目上均无统计学显著差异。学生在评估GenAI生成方案时的作业成绩显著更高,但这种局部优势并未转化为后续的总结性学习收益。调查数据进一步显示,多数学生报告干预未改变学习习惯;不过,那些报告调整学习策略的学生认为GenAI评估作业显著更有帮助。这些发现表明,GenAI评估将学生精力从开放式方案构建转向验证、诊断和判断,但不会自动产生更强的概念迁移。我们得出结论,GenAI评估活动可纳入算法课程而不会造成广泛的成绩损失,但要获得有意义的学习收益,可能需要刻意的支架式教学,推动学生超越简单的错误诊断。

英文摘要

As Generative AI (GenAI) tools become increasingly capable of generating solutions to computing assignments, the computing education community is exploring pedagogical approaches that emphasize solution evaluation, verification, and critique alongside traditional solution generation. However, evidence regarding the impact of such evaluation-centered tasks on student learning remains limited, particularly in upper-division, theory-heavy courses. We conducted a randomized A/B crossover study (N=220) in a junior-level algorithms course to compare evaluating GenAI-generated solutions with traditional problem solving. Across six assignments, student working groups either solved challenging algorithmic problems directly or evaluated often-flawed GenAI-generated solutions, with roles reversed midway through the semester. We found no statistically significant differences between groups in midterm scores, final exam scores, overall course grades, or exam problems structurally aligned with the homework interventions. Students received significantly higher homework scores when evaluating GenAI-generated solutions, but this localized advantage did not translate into downstream summative gains. Survey data further indicated that most students reported no change in study habits in response to the intervention; however, those who reported adapting their study strategies rated the GenAI-evaluation assignments as significantly more helpful. These findings suggest that GenAI evaluation redistributes student effort from open-ended solution construction toward verification, diagnosis, and judgment, but does not automatically produce stronger conceptual transfer. We conclude that GenAI-evaluation activities can be incorporated into algorithms coursework without broad performance losses, but meaningful learning gains may require deliberate scaffolding that pushes students beyond simple error diagnosis.

发表机构

  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑