跨语音与面部的情感:多模态基础模型中的共享情感机制
Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models
浏览论文内容
中文总结 AI 辅助
该研究在Gemma-4-12B-it等3款多模态基础模型中识别出情感敏感神经元,发现语音与面部情感识别的表征在解码器级部分汇聚,且存在双向因果迁移,为跨模态情感机制提供了新分析。
中文摘要 AI 辅助
现代多模态基础模型(Multimodal Foundation Models, MFMs)在涉及语音、视觉与语言整合感知的任务(包括情感识别)上取得了快速进展,但目前仍不清楚它们是通过共享情感功能单元还是模态特定通路来识别语音与面部情感。我们在三个MFMs(Gemma-4-12B-it、MiniCPM-o-4.5和Qwen2.5-Omni-7B)中探究了情感敏感神经元(Emotion-Sensitive Neurons, ESNs),即与情感类别选择性关联的稀疏解码器神经元。以语音情感识别和面部表情识别作为互补探针,我们识别出了声学ESNs和视觉ESNs。视觉ESNs具有因果意义:停用它们会选择性损害相关面部情感的识别,而调控其激活会选择性增强该情感相对于其他情感类别的识别。声学与视觉ESNs还表现出情感匹配的重叠及相似的层分布,表明语音与面部情感表征间存在部分结构对齐。最后,跨模态干预揭示了双向因果迁移:从一种模态识别出的ESNs应用于另一种模态时会产生情感特定效应。我们的研究是首批对MFMs中情感功能单元开展的跨模态激活水平分析之一,表明语音与面部情感识别部分汇聚于可定位和操控的稀疏解码器级组件,无需训练即可实现。
英文摘要
Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition. However, it remains unclear whether they recognize speech and facial emotion through shared affective functional units or modality-specific pathways. We explore emotion-sensitive neurons (ESNs), sparse decoder neurons selectively associated with emotion categories, in three MFMs: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B. Using speech emotion recognition and facial expression recognition as complementary probes, we identify acoustic and visual ESNs. Visual ESNs are causally meaningful: deactivating them selectively impairs recognition of the associated facial emotion, whereas steering their activations selectively enhances recognition of that emotion relative to other emotion categories. Acoustic and visual ESNs further show emotion-matched overlap and similar layer-wise distributions, indicating partial structural alignment between affective representations across speech and faces. Finally, cross-modal interventions reveal bidirectional causal transfer: ESNs identified from one modality produce emotion-specific effects when applied to the other. Our findings provide one of the first cross-modality activation-level analyses of affective functional units in MFMs, suggesting that speech and facial emotion recognition partially converge onto sparse decoder-level components that can be localized and manipulated without training.
发表机构
- Center for Language and Speech Processing (CLSP), Johns Hopkins University(约翰霍普金斯大学语言与语音处理中心)
- Group on Language, Audio & Music (GLAM), Imperial College London(伦敦帝国理工学院语言、音频与音乐组)
机构由 AI 辅助整理,请以论文原文为准。