发表机构
Faculty of Informatics; Masaryk University; Center for ICT Infrastructure; Yamaguchi University; Kyushu University(信息学院; 马萨里克大学; 信息与通信基础设施中心; 山阳大学; 九州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究计算教育中团队问题解决练习的评估方法,通过比较聚类和大语言模型,利用原始数据集与教师评分对比,发现聚类有效可靠、计算要求低,GPT-5.2误差低,相关方法已集成到开源平台,还共享了数据和工具。
AI 中文摘要
这篇研究实践方向的全文介绍了在桌面练习(TTXs)中评估学生团队的方法。TTXs能让学习团队为工作任务做准备并练习危机应对,如解决网络安全事件。评估对确定团队实现学习目标的程度至关重要,但TTXs复杂、开放式的性质常导致反馈延迟或不完整。TTX学习平台可记录团队行动和交流,但利用这些数据评估表现的研究不足。为填补这一空白,我们使用来自两国8位参与者的原始数据集比较了两种TTX后团队评估方法——聚类和大语言模型(LLMs)。我们根据标准化评分标准将这些方法与教师评分进行评估。聚类将以相似方式处理TTX任务的团队分组,使教师能更快地向聚类中的团队提供有针对性的反馈。这种方法有效且可靠,计算要求低。LLMs使用标准化评分标准评估团队交流。虽然GPT-4o常与教师评分不一致,但GPT-5.2的误差明显更低。研究的方法已集成到开源TTX学习平台INJECT中,以支持可扩展性和教学实践。为鼓励社区采用,我们公开共享所有数据集、软件工具和一个完整的TTX场景。
英文摘要
This full paper in the research-to-practice track presents methods for assessing student teams in tabletop exercises (TTXs). TTXs enable learner teams to prepare for workplace tasks and practice crisis responses, such as resolving cybersecurity incidents. While assessment is essential for determining how well teams achieve learning objectives, the complex, open-ended nature of TTXs often leads to delayed or incomplete feedback. TTX learning platforms can record teams' actions and communication; yet, leveraging these data to assess performance is underexplored. To address this gap, we compared two post-TTX team assessment methods -- clustering and large language models (LLMs) -- using an original dataset from 81 participants across two countries. We evaluated these methods against instructor-assigned scores based on standardized rubrics. Clustering grouped teams that approached TTX tasks similarly, enabling instructors to deliver faster, targeted feedback to teams within a cluster. This method was valid and reliable, with low computational requirements. LLMs used the standardized rubrics to assess teams' communication. While GPT-4o frequently disagreed with instructor scores, GPT-5.2 demonstrated considerably lower error. The researched methods have been integrated into INJECT, an open-source TTX learning platform, to support scalability and teaching practice. To encourage community adoption, we publicly share all datasets, software tools, and a full-fledged TTX scenario.
CommentsFull paper accepted for the main track of the IEEE FIE 2026 conference