arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06023cs.LG

BioKD:通过可靠性门实现面向情感识别的选择性生理信号到视频的知识蒸馏

BioKD: Selective Physiology-to-Video Knowledge Distillation via Reliability Gate for Emotion Recognition

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

Bojing Hou, Ruohao Li, Yitong Zhu, Hongjun Liu, Luwen Yu, Yuyang Wang

AI总结:

本文提出BioKD框架,以生理信号为训练特权信息指导视频学生模型,通过可靠性门控抑制负迁移,在DEAP、AMIGOS数据集上的情感识别任务中优于基线,且推理无额外开销。

AI中文摘要:

为解决基于视频的情感识别在模糊或社交掩盖的行为线索下存在的局限性,以及生理信号部署性差的问题,本文提出了一种感知可靠性的生理信号到视频的知识蒸馏框架,称为BioKD。该框架在训练期间利用生理信号作为特权信息,指导基于视频的学生模型学习深层情感表征,而在推理时仅依赖非侵入式的视频输入。为应对因被试间差异、信号伪影和时间不一致导致的生理教师监督存在的高噪声和不稳定性问题,BioKD结合了样本级的感知可靠性门控机制与渐进式蒸馏策略。通过自适应调节知识迁移的强度,该框架抑制了不可靠生理监督引发的负迁移,实现了更稳定的跨模态蒸馏。在DEAP和AMIGOS数据集上的实验表明,BioKD在试次级和被试级评估协议下,对于效价和唤醒度识别均始终优于代表性基线。例如,BioKD在DEAP数据集的试次级唤醒度任务上取得68.01%的性能,在更具挑战性的被试级设置下取得65.29%的性能,展现出在被试独立评估设置下的性能提升。进一步分析显示,BioKD可有效缓解教师模型的过度自信误差,且优于仅基于熵的加权策略,证实了显式建模监督可靠性的重要性。此外,BioKD相对于相同的视频学生架构未引入额外的推理时间开销,且消除了对生理感知和多模态同步的需求。

英文摘要:

To address the limitations of video-based emotion recognition under ambiguous or socially masked behavioral cues, as well as the poor deployability of physiological signals, this paper proposes a reliability-aware physiology-to-video knowledge distillation framework, termed BioKD. The proposed framework leverages physiological signals as privileged information during training to guide a video-based student model in learning deep affective representations, while relying solely on non-intrusive video inputs at inference time. To cope with the high noise and instability of physiological teacher supervision caused by inter-subject variability, signal artifacts, and temporal inconsistency, BioKD incorporates a sample-wise reliability-aware gating mechanism together with a progressive distillation strategy. By adaptively regulating the strength of knowledge transfer, the framework suppresses negative transfer induced by unreliable physiological supervision and enables more stable cross-modal distillation. Experiments on DEAP and AMIGOS show that BioKD consistently outperforms representative baselines under both trial-wise and subject-wise evaluation protocols for valence and arousal recognition. For example, BioKD achieves 68.01\% on DEAP (trial-wise arousal) and 65.29\% under the more challenging subject-wise setting, demonstrating improved performance under a subject-independent evaluation setting. Further analyses show that BioKD effectively mitigates overconfident teacher errors and outperforms an entropy-only weighting strategy, confirming the importance of explicitly modeling supervision reliability. In addition, BioKD introduces no additional inference-time overhead relative to the same video student architecture and removes the need for physiological sensing and multimodal synchronization.

↑