Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
探索多模态课堂数据中教学活动和话语的自动化识别
专题命中 其他VLM :vision-language model(abstract);分类 cs.CV
AI总结 本文通过多模态分析方法,实现了课堂活动中教学活动和话语的自动化识别,展示了微调模型在视频和 transcripts 上的高准确率,为可扩展的教师反馈系统提供了基础。
Comments This article has been accepted for publication in the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026