Versatile audio-visual learning for emotion recognition
多功能音频视觉学习用于情绪识别
机构 * Multimodal Signal Processing (MSP) Lab(多模态信号处理实验室) ; Erik Jonsson School of Engineering and Computer Science(埃里克·乔纳森工程与计算机科学学院) ; The University of Texas at Dallas(德克萨斯大学达拉斯分校)
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.MM、eess.AS
AI总结 本文提出VAVL框架,通过共享层和残差连接实现灵活的音频视觉学习,提升情绪识别和预测性能。
Comments 18 pages, 4 Figures, 3 tables (published at IEEE Transactions on Affective Computing)
Journal ref IEEE Transactions on Affective Computing 2025