在Token空间中读取情感:面向情感识别的SpeechLLM判别式适配
Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition
浏览论文内容
中文总结 AI 辅助
针对SpeechLLM生成式解码不适合情感分类的问题,提出用分类头读取最终Token隐藏状态的判别式适配,在IEMOCAP上提升Macro F1并消除幻觉,且保持可解释性。
中文摘要 AI 辅助
SpeechLLM在情感识别方面展现出强大的潜力,然而它们通过生成式解码器读取预测的情感,这种方式并不适合分类任务:它可能输出目标集之外的标签,并且偏向于频繁出现的类别。我们提出了一种判别式适配方法,通过分类头读取最终提示词Token的隐藏状态,在一次前向传播中产生标签,且无需修改骨干网络。由于这种读取方式起始于模型原本会解码的隐藏状态,因此在完全相同的SpeechLLM中,它提供了生成式推理与判别式推理的受控比较。我们将分类头保持为单层线性层,以微小的精度换取可解释性:每个情感在LLM输出Token空间中成为一个方向,从而揭示相关联的Token。在IEMOCAP上,跨两种SpeechLLM架构,该方法提升了Macro F1分数并消除了幻觉,在真实的ASR转录上取得了最大的提升。我们的分析表明,这些情感方向编码了间接关联,反映了网络规模文本中的偏见。
英文摘要
SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set and favors frequent classes. We propose a discriminative adaptation that reads the final prompt token's hidden state through a classification head, producing a label in one forward pass without modifying the backbone. Because this readout starts from the hidden state the model would otherwise decode, it gives a controlled comparison of generative and discriminative inference in an otherwise identical speechLLM. We keep the head a single linear layer, trading little accuracy for interpretability: each emotion becomes one direction in the LLM output token space, revealing associated tokens. On IEMOCAP, across two speechLLM architectures, it improves Macro F1 and removes hallucinations, with largest gains on realistic ASR transcripts. Our analysis reveals that these emotion directions encode indirect associations mirroring biases in web-scale text.
发表机构
- Idiap Research Institute(伊迪亚普研究所)
- EPFL(瑞士洛桑联邦理工学院)
- Uniphore(Uniphore公司)
- Brno University of Technology(布尔诺理工大学)
机构由 AI 辅助整理,请以论文原文为准。