arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型音频语言模型中的多语言情感神经元

Multilingual Emotion Neurons in Large Audio-Language Models

Xiutian Zhao, Philipp Koehn, Björn Schuller, Berrak Sisman

arXiv 2608.08772首次发表:更新:

发表机构

Johns Hopkins University; Imperial College London(约翰斯·霍普金斯大学; 伦敦帝国学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过首个神经元层面可解释性研究,提出多语言情感神经元(MLENs)定义及一致性正则化融合(CR-Fusion)方法,证实其在零样本和低资源场景下情感控制更精准,为LALMs跨语言情感编码提供因果解释。

AI 中文摘要

情感是人类交流的核心,其表达因语言而异。大型音频语言模型(LALMs)在多语言语音任务上表现出色,但目前尚不清楚它们是通过语言特定的关联还是语言无关的表征来编码情感。我们针对该问题开展了首个神经元层面的可解释性研究。我们将多语言情感神经元(MLENs)定义为在不同语言间表现出稳定情感选择性和一致因果效应的功能单元,并引入一致性正则化融合(CR-Fusion)方法来识别它们。在四个现代LALMs和12种类型多样的语言中,按每种语言独立识别的情感敏感神经元重叠度极低,额外的单语言识别数据会迅速饱和,无法分离出更具可迁移性的单元,这促使我们从跨语言汇总证据中进行识别。因果干预实验表明,由CR-Fusion识别的MLENs在零样本和低资源场景下,比单语言神经元集提供更精准、可迁移的情感控制。留一法消融实验进一步揭示了不对称迁移:包括低资源语言在内的各识别语言均贡献非冗余证据,而部分低资源语言从该跨语言迁移中获益最多。综上,我们的发现提供了首个关于LALMs如何跨语言编码情感的因果神经元层面解释,并确立多语言神经元识别是理解跨语言情感行为的有效机制。

英文摘要

Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks, yet it remains unclear whether they encode emotion through language-specific correlations or language-agnostic representations. We present the first neuron-level interpretability study of this question. We define Multilingual Emotion Neurons (MLENs) as functional units exhibiting stable emotional selectivity and aligned causal effects across languages, and introduce Consistency-Regularized Fusion (CR-Fusion) to identify them. Across four modern LALMs and 12 typologically diverse languages, emotion-sensitive neurons identified independently per language show minimal overlap, and additional monolingual identification data saturates quickly without isolating more transferable units, motivating identification from pooled cross-lingual evidence. Causal interventions demonstrate that MLENs identified by CR-Fusion provide more precise and transferable affective control than monolingual neuron sets in both zero-shot and low-resource settings. Leave-one-out ablations further reveal asymmetric transfer: individual identification languages, including low-resource ones, contribute non-redundant evidence, while several low-resource languages benefit most from the resulting cross-lingual transfer. Together, our findings provide the first causal, neuron-level account of how LALMs encode emotion across languages, and establish multilingual neuron identification as an effective mechanism for understanding cross-lingual affective behavior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑