arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05806cs.AI

揭示对话情感识别中的弱点

Exposing Weaknesses in Emotion Recognition in Conversations

Amir Ben Khalifa, Fanny Bezancon, Bessam Abdulrazak, Amine Trabelsi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过零样本LLM实验和人工重标注揭示对话情感识别中单标签假设的缺陷,并提出LLM-as-Judge框架按情感合理性独立评估以改进评测。

中文摘要 AI 辅助

对话情感识别(ERC)旨在识别多轮对话中说话者的情感。准确的情感识别可以支持广泛的应用,包括共情对话代理、心理健康支持和教育技术。虽然许多近期方法依赖于任务特定的微调,但此类模型可能利用数据集特定的线索。ERC中一个核心但很少被质疑的假设是,每个话语可以被分配一个单一且无歧义的情感标签。为了调查这一假设,我们在零样本设置下使用大语言模型(LLMs)研究ERC,并将前序对话轮次作为上下文纳入。我们表明,聚合指标掩盖了系统性失败。错误集中在包含否定、感叹和插入语的话语周围。这一模式在所有评估的模型中一致出现,表明是基准测试的局限性而非模型特定的弱点。一项涉及四位人工标注者的受控重新标注研究支持了这一发现:仅在35%的案例中观察到强一致性,其中中性话语主导了高一致性实例,而许多情感类别落入低一致性区域。这些发现表明,许多表面上的模型错误反映了真实的标注歧义,而非情感理解能力差。因此,标准的单标签评估是不充分的。为解决这一局限性,我们引入了一个LLM作为评判者的框架,该框架根据每种情感在对话语境中的合理性独立评估,而非强制执行单一标签决策。

英文摘要

Emotion Recognition in Conversations (ERC) aims to identify speakers' emotions in multi-turn dialogue. Accurate emotion recognition can support a wide range of applications, including empathetic conversational agents, mental health support, and educational technologies. While many recent approaches rely on task-specific fine-tuning, such models may exploit dataset-specific cues. A central yet rarely questioned assumption in ERC is that each utterance can be assigned a single unambiguous emotion label. To investigate this assumption, we study ERC using Large Language Models (LLMs) in a zero-shot setting while incorporating preceding conversational turns as context. We show that aggregate metrics mask systematic failures. Errors concentrate around utterances containing negations, exclamations, and interjections. This pattern is consistent across all evaluated models, suggesting limitations in the benchmarks rather than model-specific weaknesses. A controlled re-annotation study involving four human annotators supports this finding: strong agreement is observed in only 35 percent of cases, with neutral utterances dominating high-agreement instances, while many emotional categories fall into low-agreement regimes. These findings suggest that many apparent model errors reflect genuine annotation ambiguity rather than poor emotion understanding. Standard single-label evaluation is therefore insufficient. To address this limitation, we introduce an LLM-as-Judge framework that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision.

发表机构

  • Université de Sherbrooke(谢布鲁克大学)
  • Université de Nantes(南特大学)

机构由 AI 辅助整理,请以论文原文为准。

↑