发表机构
Emory University(埃默里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于LLM的多模态对话维度情感评估框架,通过LoRA微调LLaMA模型在IEMOCAP上实现Valence CCC 0.7822的新最先进水平,并发现音频线索对小型模型增益显著。
AI 中文摘要
对话中的情感识别已被广泛研究,但将大语言模型(LLMs)应用于多模态对话中的连续维度情感评估仍 largely unexplored。我们提出一个基于LLM的框架,在IEMOCAP数据集上执行离散情感识别和Valence-Arousal-Dominance(VAD)维度评估,该框架遵循SpeechCueLLM方法,将声学线索作为自然语言描述融入。我们评估了涵盖LLaMA、GPT和Qwen系列的六个模型,在零样本提示、少样本提示和LoRA微调设置下进行测试。LoRA微调的LLaMA模型在两项任务上均显著优于提示工程的GPT模型,尽管GPT规模更大,我们将这一差距归因于领域适应而非模型能力。我们的最佳模型在Valence上达到CCC 0.7822,创下IEMOCAP上的新最先进水平。消融研究证实,文本音频描述能有效提升较小模型(加权F1提高3.5至3.6),但对最大模型贡献甚微,表明当语言能力有限时,音频线索最有价值。VAD维度上的性能不对称性与IEMOCAP自身注释中的标注者一致性层级密切相关。
英文摘要
Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensional evaluation on IEMOCAP, incorporating acoustic cues as natural language descriptions following the SpeechCueLLM approach. We evaluate six models spanning the LLaMA, GPT, and Qwen families under zero-shot prompting, few-shot prompting, and LoRA fine-tuning. LoRA fine-tuned LLaMA models substantially outperform prompt-engineered GPT models on both tasks despite GPT's larger scale, a gap we attribute to domain adaptation rather than model capacity. Our best model achieves a Valence CCC of 0.7822, a new state-of-the-art on IEMOCAP. Ablation studies confirm that textual audio descriptions meaningfully improve smaller models (+3.5 to 3.6 weighted F1) while contributing little for the largest model, suggesting audio cues are most valuable when linguistic capacity is limited. The performance asymmetry across VAD dimensions closely mirrors the annotator agreement hierarchy in IEMOCAP's own annotations.
Comments15 pages, 6 figures, 11 tables