课堂视频多任务情感状态识别中的维度特异性不平衡与自适应混合标签策略
Dimension-Specific Imbalance and an Adaptive Hybrid Label Strategy for Multi-Task Affective State Recognition in Classroom Video
浏览论文内容
中文总结 AI 辅助
针对课堂视频情感识别中的维度特异性类别不平衡问题,提出自适应混合标签策略(AHLS),结合四级/二元分类与空类占位符,在DAiSEE上以轻量模型取得最优平均Macro-F1 47.70%。
中文摘要 AI 辅助
从课堂视频中识别学生情感状态受到类别不平衡问题的制约,我们认为该问题的本质尚未得到充分分析。我们证明,在DAiSEE(该任务的事实标准基准)上,类别不平衡从根本上是一种维度特异性现象:原始不平衡比率最高可达约79:1(挫败感),且不平衡的结构(不仅仅是其幅度)在不同维度间存在差异,因此统一的算法处理(激进的类别重加权、焦点损失、统一二元简化)会以依赖维度的方式失效。我们提出了自适应混合标签策略(AHLS),该策略为不平衡程度可控的维度分配四级分类,为严重长尾分布的维度分配二元分类,并结合空类占位符机制,通过将四个维度上的最大有效逆频率权重比从高达约79倍压缩至至多约8.5倍,来稳定多任务优化。在轻量级FERShuffleNetV2 + LTCN架构(0.39百万参数,0.07 G FLOPs)上验证,该策略在11种对比方法中,包括五种最先进的深度模型(最大大167倍)和五种传统机器学习基线,在所有四个情感维度上取得了最高的平均Macro-F1,为47.70%。在DAiSEE上的跨模型行为分析表明,在此设置下,分类策略可能比单纯扩大骨干网络规模对少数状态识别的影响更强。我们还报告了DAiSEE在少数情感状态上存在标注密度限制的证据,这表明基准设计可能像算法改进一样限制未来的进展。索引术语:情感计算、课堂视频分析、类别不平衡、多任务学习、轻量级深度学习、DAiSEE、参与度识别。
英文摘要
Recognition of student affective states from classroom video is constrained by a class imbalance problem whose true nature is, we argue, under-analyzed. We show that on DAiSEE -- the de facto benchmark for this task -- class imbalance is fundamentally a dimension-specific phenomenon: raw imbalance ratios reach as high as approximately 79:1 (Frustration), and the structure of imbalance -- not only its magnitude -- varies across dimensions, so uniform algorithmic treatments (aggressive class reweighting, focal loss, uniform binary simplification) fail in dimension-dependent ways. We propose the Adaptive Hybrid Label Strategy (AHLS), which assigns four-level classification to dimensions with manageable imbalance and binary classification to severely long-tailed dimensions, coupled with a null-class placeholder mechanism that stabilizes multi-task optimization by compressing the maximum effective inverse-frequency weight ratio from as high as ~79$\times$ to at most ~8.5$\times$ across the four dimensions. Validated on a lightweight FERShuffleNetV2 + LTCN architecture (0.39 M parameters, 0.07 G FLOPs), the strategy attains the highest mean Macro-F1 of 47.70% across all four affective dimensions among 11 compared methods, including five state-of-the-art deep models (up to 167$\times$ larger) and five traditional machine-learning baselines. A cross-model behavioral analysis on DAiSEE indicates that, in this setting, classification strategy may influence minority-state recognition more strongly than further backbone scaling alone. We also report evidence of an annotation-density limitation of DAiSEE on minority affective states, which suggests that benchmark design may bound future progress as much as algorithmic refinement. Index Terms Affective computing, classroom video analysis, class imbalance, multi-task learning, lightweight deep learning, DAiSEE, engagement recognition.
发表机构
- The Education University of Hong Kong(香港教育大学)
机构由 AI 辅助整理,请以论文原文为准。