发表机构
The Education University of Hong Kong; Yew Chung College of Early Childhood Education(香港教育大学; 耀中幼教学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较GPT-5模型与人类评分员在课堂观察中的CLASS评分,发现AI在情感支持领域趋同但整体无法替代人类,更适合作为初步筛查工具。
AI 中文摘要
课堂观察被广泛认为是建立教育质量基准和指导教学改进的关键工具,然而它们仍然资源密集且依赖于训练有素的观察者。本研究评估了使用大型语言模型(GPT-5模型)对幼儿课堂中师幼互动进行评分的可行性,并以人类评分员作为基准。研究分析了来自香港30所幼儿园38个班级的87个视频录制观察。利用观察转录文本,AI模型被配置为应用完整的课堂评估评分系统(CLASS)框架。随后,通过检查CLASS领域和维度的相关性及平均分差异,将AI评分与人类评分进行比较。结果显示,在情感支持领域,尤其是反馈质量维度(该维度捕捉教师如何使用反馈来拓展儿童学习),AI与评分员之间的趋同性更高。而在更具程序性或依赖情境的互动中,特别是在课堂组织和教学支持领域,出现了更大的分歧。这些发现表明,基于转录文本的AI评分可能捕捉到师幼互动中的部分相对差异,但尚不能始终如一地复现经过校准的人类判断以覆盖完整的CLASS框架。因此,AI辅助观察可能更适合作为初步筛查工具,而非替代训练有素的观察者,为教师提供用于反思的证据,而非用于高风险评估。未来研究应探讨领域特定训练以及结合情境和视觉信息是否能改善AI与人类评分之间的一致性。
英文摘要
Classroom observations are widely recognized as a key tool for establishing benchmarks of education quality and guiding pedagogical improvement, yet they remain resource-intensive and dependent on trained observers. This study evaluated the feasibility of using a LLM (GPT-5 model) to score teacher-child interactions in early childhood classrooms, benchmarked against human raters. The study analyzed 87 video-recorded observations from 38 classrooms across 30 kindergartens in Hong Kong. Using observation transcripts, the AI model was configured to apply the full Classroom Assessment Scoring System (CLASS) framework. AI-rated scores were then compared with human ratings by examining correlations and differences in mean scores of the CLASS domains and dimensions. The results showed greater convergence between AI and raters for the Emotional Support domain and, in particular, the Quality of Feedback dimension, which captures how teachers use feedback to extend children's learning. Greater divergence emerged for interactions that were more procedural or context-dependent, particularly within the Classroom Organization and Instructional Support domains. These findings suggest that transcript-based AI scoring may capture some of the relative variation in teacher-child interactions but cannot yet reproduce calibrated human judgements consistently across the full CLASS framework. AI-assisted observation may therefore be more appropriate as a preliminary screening tool rather than as a replacement for trained observers, providing teachers with evidence for reflection rather than high-stakes evaluation. Future research should examine whether domain-specific training and incorporation of contextual and visual information can improve alignment between AI and human rated scores.