arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16318cs.SEcs.AIcs.PF

重新审视生成式人工智能在面向对象编程入门评估中的表现:来自2026年的见解

Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming Assessments: Insights from 2026

Marina Lepp, Joosep Kaimre

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估ChatGPT-5.2等5种GenAI系统在OOP入门课程评估中的表现,发现其得分高于平均学生,在部分任务有优势但仍存在编译、高级OOP概念等问题,且较前一年有明显改进。

中文摘要 AI 辅助

生成式人工智能(GenAI)的最新进展大幅提升了大型语言模型(LLM)生成和解释源代码的能力。然而,它们在真实的面向对象编程(OOP)评估中的表现仍未得到充分理解。本研究使用某大学OOP入门课程的编程测试和考试任务,评估了五种广泛使用的GenAI系统:ChatGPT-5.2、DeepSeek-V3、Gemini 2.5 Flash、Claude Sonnet 4.5和M365 Copilot。生成的解决方案采用与学生相同的评分标准进行评估,并与同一课程的历史学生成绩以及前一年的研究结果进行比较。还分析了常见错误,以识别各模型反复出现的局限性。所有被评估的GenAI系统的得分均高于平均学生群体,且在较长的编程任务中经常获得满分。不过,它们偶尔会生成无法编译的代码,并且在高级OOP概念方面仍然存在困难,尤其是接口、抽象类和某些与继承相关的任务。在涉及图像解释的图形相关问题上,性能也受到限制。与前一年相比,被评估的系统在大多数评估中表现出明显的改进,同时呈现出几种反复出现的错误模式。这些发现提供了当代GenAI系统在真实OOP入门评估中的能力和局限性的最新评估,也为编程评估的设计、GenAI工具在软件工程教育中的负责任整合以及未来评估AI辅助编程演变的研究提供了可参考的证据。

英文摘要

Recent advances in Generative Artificial Intelligence (GenAI) have substantially improved the ability of large language models (LLMs) to generate and explain source code. However, their performance on authentic object-oriented programming (OOP) assessments remains insufficiently understood. This study evaluates five widely used GenAI systems, ChatGPT-5.2, DeepSeek-V3, Gemini 2.5 Flash, Claude Sonnet 4.5, and M365 Copilot, using programming tests and examination tasks from an introductory university OOP course. The generated solutions were assessed using the same grading criteria applied to students and compared with historical student results from the same course, as well as findings from the previous year. Common errors were also analyzed to identify recurring limitations across models. All evaluated GenAI systems achieved higher scores than the average student cohort and frequently obtained full marks on longer programming tasks. Nevertheless, they occasionally produced non-compiling code and continued to struggle with advanced OOP concepts, particularly interfaces, abstract classes, and certain inheritance-related tasks. Performance was also limited on graphics-related questions involving image interpretation. Compared with the previous year, the evaluated systems demonstrated noticeable improvements across most assessments while exhibiting several recurring error patterns. The findings provide an updated evaluation of the capabilities and limitations of contemporary GenAI systems on authentic introductory OOP assessments. They also offer evidence that can inform the design of programming assessments, the responsible integration of GenAI tools into software engineering education, and future studies evaluating the evolution of AI-assisted programming.

发表机构

  • University of Tartu(塔尔图大学)

机构由 AI 辅助整理,请以论文原文为准。

↑