arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

正确性、收敛性与AI生成代码检测:入门编程中学生与大型语言模型代码的纵向研究

Correctness, Convergence, and AI-Generated Code Detection: A Longitudinal Study of Student and Large Language Model Code in Introductory Programming

Runlong Ye, Jing Fan, Angela Zavaleta Bernuy, Oscar Karnalim, Paul Denny, Juho Leinonen, Michael Liut

arXiv 2610.00863首次发表:更新:

发表机构

University of Toronto; Aalto University; McMaster University; Maranatha Christian University; University of Auckland; University of Toronto Mississauga(多伦多大学; 阿尔托大学; 麦克马斯特大学; 马拉纳撒基督教大学; 奥克兰大学; 多伦多大学密西沙加分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过纵向分析学生与LLM代码,发现生成参考匹配可追踪群体代码趋同,但归因个体AI使用需额外证据。

AI 中文摘要

大型语言模型能够为编程作业生成看似合理的解决方案,这使得通过将学生代码与生成的参考答案库进行匹配来检测其使用变得颇具吸引力。然而,当作业只允许少数几种自然实现时,类似的代码也可能出现,这便引发了一个问题:匹配结果究竟说明了什么。我们利用2021年、2023年和2025年开设的十个Python实验的29,970份学生提交,以及三个前沿大型语言模型回顾性生成的90,000次解决方案尝试,对生成参考匹配进行了研究。我们使用隐藏的教师测试验证生成的解决方案,在排除起始代码后使用MOSS比较代码,并检查所选函数的精确抽象语法树(AST)形式。这些模型通常能生成正确的解决方案,并且在大多数作业中,它们收敛于相似的实现。在较晚的队列中,学生提交与生成的参考匹配更为频繁,包括那些通过所有隐藏测试的提交。对于规定严格的函数,模型收敛于少数几种精确的抽象语法树形式,且不同学生形式的数量在各队列中也呈下降趋势,而开放式函数在两种来源中均保持多样性。大多数报告的匹配重叠部分较短,这使得最小匹配长度成为审查学生代码时的一个重要选择。最后,我们讨论了教师如何在发布作业之前构建生成解决方案的参考库,以识别生成解决方案收敛的任务,决定匹配应引起多大程度的审查,并重新设计任务以引出测试、推理和中间工作。这些发现支持跟踪提交代码的群体层面变化,而将AI使用归因于单个提交则需要关于其产生方式的额外证据,如提示词、修订、中间代码和学生的披露。

英文摘要

Large language models can generate plausible solutions to programming assignments, making it tempting to detect their use by matching student code against a reference bank of generated solutions. Yet similar code can also arise when an assignment admits only a few natural implementations, which leaves open what a match actually shows. We investigate generated-reference matching using 29,970 student submissions from ten Python labs offered in 2021, 2023, and 2025, together with 90,000 solution attempts generated retrospectively by three frontier LLMs. We validate the generated solutions using hidden instructor tests, compare code with MOSS after excluding the starter code, and examine the exact abstract syntax tree (AST) forms of selected functions. The models usually produced correct solutions and, across most assignments, converged on similar implementations. Student submissions matched the generated references more often in later cohorts, including among submissions that passed every hidden test. On tightly specified functions, the models converged on a few exact abstract-syntax-tree forms, and the number of distinct student forms also declined across cohorts, whereas open-ended functions remained diverse in both sources. Most reported overlaps were short, making the minimum match length an important choice when reviewing students' code. Finally, we discuss how instructors can build a reference bank of generated solutions before releasing an assignment to identify tasks on which generated solutions converge, decide how much review a match warrants, and redesign tasks to elicit tests, reasoning, and intermediate work. These findings support tracking population-level changes in submitted code, while attributing AI use to an individual submission would require additional evidence about how it was produced, such as prompts, revisions, intermediate code, and student disclosures.

Commentsaccepted at Koli Calling '26

DOI:10.1145/3856208.3856226

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑