发表机构
The Fu Foundation School of Engineering and Applied Science, Columbia University; School of Nursing, Columbia University(哥伦比亚大学富兰克林工程与应用科学学院; 哥伦比亚大学护理学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LALMs在语调与字面意义不一致场景下的表现问题,构建CREMA-ASIS数据集,经实验发现其受语义主导且难识别不一致,监督式微调可显著提升其性能。
AI 中文摘要
跨模态的情感线索可能存在不一致(例如讽刺或嘲讽式赞美),仅依赖单一模态可能导致误解。大型音频语言模型(Large Audio-Language Models, LALMs)近来颇受欢迎,并被应用于多模态情感识别,但其解耦声学与语义线索的能力,尤其是在不一致场景下的表现,仍未得到充分探索。为解决这一空白,我们引入CREMA-ASIS数据集,该数据集专门用于研究声学情感与语义情感线索间的不一致性,它将声学情感标签与语义情感极性配对。利用该数据集,我们在多任务框架内评估LALMs的偏差,并开展分层分析以识别各层的模态主导性。我们的发现显示,LALMs在语义-声学不一致场景中表现不佳,极少预测到不一致情况,且主要受语义信息影响。不过,监督式微调可显著提升LALMs在CREMA-ASIS测试集上的性能,同时保持转录准确性与联合情感识别能力。研究结果表明,该方法具备提升LALMs在域外数据上的声学与语义理解能力的潜力。
英文摘要
Affective cues across modalities may be incongruous (e.g., sarcasm or mocking praise), potentially leading to misinterpretation when relying on a single modality. Large Audio-Language Models (LALMs) have recently gained popularity and been applied to multimodal emotion recognition, but their ability to disentangle acoustic and semantic cues, especially in incongruent cases, remains underexplored. To address this gap, we introduce CREMA-ASIS, a dataset specifically created to investigate incongruence between acoustic emotion and semantic sentiment cues. It pairs acoustic emotion labels with semantic sentiment polarities. Using this dataset, we evaluate LALM biases within a multitask framework and conduct a layer-wise analysis to identify modality dominance across layers. Our findings reveal that LALMs struggle with semantic-acoustic incongruent cases, rarely predicting incongruity, and that LALMs are predominantly influenced by semantic information. However, supervised fine-tuning significantly improves LALM performance on our CREMA-ASIS test set while preserving transcription accuracy and joint emotion recognition. Results demonstrate potential for enhancing both acoustic and semantic understanding on out-of-domain data.
CommentsEMNLP 2026 Findings