arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成式人工智能(GenAI)时代的评估设计:用于测试学生AI素养、学习成果与反思的X1-X2-X3评估模式

Assessment Design in the GenAI Era: The X1-X2-X3 Assessment Pattern for Testing Students' AI Literacy, Learning Outcomes, and Reflection

Riasat Islam, Thomas Roelleke

arXiv 2608.12351首次发表:更新:

发表机构

Queen Mary University of London(伦敦玛丽女王大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对GenAI对在线评估的挑战,提出可复用的X1-X2-X3评估模式,用于测试学生AI素养等,为教师适配GenAI时代评估提供实践指导。

AI 中文摘要

生成式人工智能(GenAI)对无监督在线评估的有效性构成了挑战,尤其是在技术学科中,只需付出少量努力就能生成看似合理的答案。本文报告了在一门大型二年级本科数据库系统课程中,设计并实施一种感知AI、测试AI的评估方法所获得的经验。该设计结合了两个关联要素:(1)结构化的三部分作答格式(X1-X2-X3),要求学生记录来源答案、生成自身答案并评估来源输出;(2)感知AI的问题设计流程,其中对任务草案进行当代GenAI工具的压力测试,当通用提示生成的答案表面上足够时则进行修订。该论述基于存档的评估材料、评分标准、规划记录、设计时的GenAI测试、练习作答数据、成绩记录及外部评审意见。其主要贡献是一种可复用的评估设计方法,而非对可测量学习成果的主张。本文展示了该模式如何在迭代中发展,以及它如何支持真实性评估、可见的AI素养、学生判断和更透明的评分。本文为教师调整评估以适应GenAI的日常使用提供了实践指导,重点在于测试AI素养而非因学生不当行为进行惩罚。

英文摘要

Generative artificial intelligence (GenAI) has challenged the validity of unsupervised online assessment, especially in technical subjects where plausible answers can be produced with little effort. This paper reports lessons from designing and implementing an AI-aware, AI-testing assessment in a large second-year undergraduate database systems module. The design combined two linked elements: (1) a structured three-part response format (X1-X2-X3) in which students documented a sourced answer, produced their own answer, and evaluated the sourced output; and (2) an AI-aware question-design process in which draft tasks were stress-tested against contemporary GenAI tools and revised when generic prompting produced superficially adequate answers. The account draws on archived assessment materials, rubrics, planning records, design-time GenAI trials, practice-response data, attainment records, and external review comments. Its main contribution is a reusable assessment-design method rather than a claim of measured learning gains. We show how the pattern developed across iterations and how it can support authentic assessment, visible AI literacy, student judgement, and more transparent marking. The paper offers practical guidance for lecturers adapting assessment to routine GenAI use, focusing on testing AI literacy rather than penalising students for misconduct.

Comments30 pages, 1 figure, 11 tables. Submitted to the following journal: Assessment & Evaluation in Higher Education (Taylor & Francis)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑