ChatGPT 解决所有经测试的 Qiskit 作业
ChatGPT Solves All Tested Qiskit Homework Assignments
浏览论文内容
中文总结 AI 辅助
本研究测试的个性化、面向执行的 Qiskit 作业设计,无法阻止 ChatGPT 成功完成,各测试实例的 ChatGPT 抗性为零,需辅以直接评估保障教育效果。
中文摘要 AI 辅助
生成式 AI 给量子软件教育带来了评估挑战:学生可以将作业笔记本提供给 ChatGPT 并要求完成提交。本研究探讨能否在要求学生运行、审查和讨论结果而非禁止 AI 的前提下,保持入门级 Qiskit 作业的可自动评分性。测试了三个作业包:带比特翻转的种子基态电路与自定义测量映射、量子傅里叶变换及其逆变换恢复、带自定义预言机掩码的种子 Deutsch-Jozsa。这些设计采用了个性化、模拟器执行、JSON 提交、隐藏引用、电路指标、反思内容及可选的 IBM Quantum 执行。每个作业包各有一个学生可见实例,在 50 次独立 ChatGPT 会话中测试,共 150 次会话。所有最终工件均已执行并通过评分。9 次会话被完整归档;均无需修改算子代码或修正量子逻辑。根据本研究的操作定义,每个测试实例的 ChatGPT 抗性均为零。种子改变的是参数而非任务结构,预期结果可从可见作业逻辑推导,脚手架暴露关键解题步骤,隐藏评分验证输出一致性,无需确立独立作者身份或理解。由于每个作业包仅重复一个实例,结果未证明每个种子或所有 Qiskit 评估均可解。因此,测试的个性化、面向执行的带回家作业设计,无法阻止学生以最低参与度的工作流程成功完成。正确的工件应辅以通过监督修改、口头答辩、预测及迁移任务进行的直接评估。
英文摘要
Generative AI creates an assessment challenge in quantum software education: a student can provide a homework notebook to ChatGPT and request a completed submission. This study examined whether introductory Qiskit homework could remain autogradable while requiring students to run, review, and discuss results rather than banning AI. Three packages were tested: seeded basis-state circuits with bit flips and customized measurement mappings; Quantum Fourier Transform followed by inverse-transform recovery; and seeded Deutsch-Jozsa with customized oracle masks. The designs used personalization, simulator execution, JSON submissions, hidden references, circuit metrics, reflections, and optional IBM Quantum execution. For each package, one student-visible instance was tested in 50 separate ChatGPT sessions, yielding 150 sessions overall. Every final artifact was executed and passed its grader. Nine sessions were fully archived; none required operator code changes or correction of quantum logic. Under the study's operational definition, each tested instance had zero observed ChatGPT-resiliency. Seeds changed parameters rather than task structure, expected results remained derivable from visible assignment logic, scaffolding exposed key solution steps, and hidden grading verified output consistency without establishing independent authorship or understanding. Because one instance was repeated for each package, the results do not establish solvability for every seed or possible Qiskit assessment. The tested personalized, execution-oriented take-home designs therefore did not prevent successful completion under a minimally engaged-student workflow. Correct artifacts should be complemented by direct assessment through supervised modification, oral defense, prediction, and transfer tasks.