MUPA²E:用于情绪评估的非对称注意力多模态统一感知框架
MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment
浏览论文内容
中文总结 AI 辅助
本文提出MUPA²E多模态统一感知框架,用共享非对称注意力骨干处理面部视频与EEG,在DMER数据集上对比多配置,发现时长线索影响分类,控制时长后准确率下降,证明统一架构处理异质信号的可行性。
中文摘要 AI 辅助
自动情绪评估可从神经信号与行为信号的结合中获益,但许多多模态方法在融合前依赖独立的、针对特定模态的特征提取流水线。本文提出MUPA²E,这是一个统一感知框架,通过单一共享的非对称注意力骨干网络处理面部视频和脑电图(EEG)。面部视频通过轴折叠帧令牌表示,而EEG可作为原始多通道波形处理,或投影到空间域以进行多模态融合。该框架在DMER数据集上采用分层的受试者独立协议进行评估,比较单模态视频、单模态EEG以及融合视频-EEG配置(含逐通道和合并EEG投影)。使用原始记录,将较短试零填充以匹配最长时长,步长为30的合并融合达到最高验证性能,测试准确率为70.07%。进一步分析发现,情感类别间的记录时长分布不均,使得填充模式成为潜在分类线索。通过将所有记录裁剪为20秒的共同时长来控制该因素,测试准确率为62.71%,为该框架提供了更严格的时长控制评估,消除了记录长度差异作为潜在分类线索。这些发现证明了在紧凑的统一架构内处理结构不同的神经和视觉信号的可行性,同时强调了在情感数据集中控制时长相关线索的重要性。
英文摘要
Automatic emotion assessment can benefit from combining neural and behavioral signals, but many multimodal approaches rely on separate, modality-specific feature-extraction pipelines before fusion. This paper presents MUPA\textsuperscript{2}E, a unified perception framework that processes facial video and electroencephalography (EEG) through a single shared asymmetric-attention backbone. Facial video is represented through axis-folded frame tokens, while EEG is processed either as a raw multichannel waveform or projected into the spatial domain for multimodal fusion. The framework is evaluated on the DMER dataset under a stratified subject-independent protocol, comparing unimodal video, unimodal EEG, and fused video--EEG configurations with per-channel and merged EEG projections. Using the original recordings, with shorter trials zero-padded to match the longest duration, merged fusion at stride~$30$ achieves the highest validation performance and a test accuracy of $70.07\%$. Further analysis revealed that recording duration is unevenly distributed across the affective classes, making the padding pattern a potential classification cue. Controlling for this factor by cropping all recordings to a common duration of $20$ seconds yielded a test accuracy of $62.71\%$, providing a stricter duration-controlled assessment of the framework in which differences in recording length are removed as a potential classification cue. These findings demonstrate the feasibility of processing structurally different neural and visual signals within a compact unified architecture while highlighting the importance of controlling duration-related cues in affective datasets.
发表机构
- Honda Research Institute Japan(日本本田研究所)
机构由 AI 辅助整理,请以论文原文为准。