发表机构
University of Amsterdam(阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在DIPSER数据集上评估公平感知多模态时间模型,发现多模态融合可适度提升学生注意力估计的预测与最差组性能,但验证级公平增益未必可推广,需采用子组感知评估等方法开展稳健公平性评估。
AI 中文摘要
自动学生注意力估计可支持学习分析,但聚合预测指标可能掩盖人口统计学差异。本研究在DIPSER(结合面部图像、可穿戴传感器测量数据、注意力标注及自动推断人口统计元数据的自然课堂数据集)上评估公平感知多模态时间模型。在10个训练随机种子下,对比三种基线模型:视觉GRU、传感器GRU及残差融合Transformer。该多模态模型取得最佳平均测试性能(MAE 0.283,RMSE 0.363),且在评估的基线中具有最低的最差组误差,尽管其相较于视觉GRU的提升较为有限。针对性别与年龄的MAE-gap正则化可减少验证数据上的差异,但此类增益无法一致迁移至保留的被试或重复的被试级划分。在NVIDIA A100-SXM4-40GB GPU上,预热后的端到端管道以1秒步长平均每个预测窗口耗时50.65毫秒,而时间模型本身仅需1.02毫秒。研究结果表明,多模态融合可适度提升预测及最差组性能,但不应假设验证级别的公平性增益具有可推广性,因此,稳健的公平性评估需采用子组感知评估、重复的被试级验证及更大、人口统计学样本更均衡的数据集。
英文摘要
Automated student-attention estimation can support learning analytics, but aggregate predictive metrics can conceal demographic disparities. This study evaluates fairness-aware multimodal temporal models on DIPSER, a naturalistic classroom dataset combining facial images, wearable-sensor measurements, attention annotations, and automatically inferred demographic metadata. Three baselines are compared across 10 training seeds: a Visual GRU, a Sensor GRU, and a Residual Fusion Transformer. The multimodal model achieves the best mean test performance (MAE 0.283, RMSE 0.363) and the lowest worst-group error among the evaluated baselines, although its gain over the Visual GRU is modest. Gender- and age-targeted MAE-gap regularization reduces disparities on validation data, but these gains do not consistently transfer to held-out subjects or repeated subject-level splits. On an NVIDIA A100-SXM4-40GB GPU, the warm end-to-end pipeline averages 50.65 ms per prediction window at a one-second stride, while the temporal model itself requires 1.02 ms. The findings show that multimodal fusion can modestly improve prediction and worst-group performance, but validation-level fairness gains should not be assumed to generalize. Robust fairness assessment therefore requires subgroup-aware evaluation, repeated subject-level validation, and larger, better balanced demographic samples.