发表机构
Tianjin University; Fuzhou University; Shanghai Jiaotong University; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(天津大学; 福州大学; 上海交通大学; 中国科学院深圳先进技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对口语语言模型情商评估缺乏理论框架的问题,构建了EmoSBench基准,开发了基于SFT和GRPO优化的EmoS评估模型,其在EmoSBench上准确率达83.8%,接近人类水平。
AI 中文摘要
尽管遵循指令和听觉理解方面已取得显著进展,但口语语言模型(SLM)的情商(EI)评估仍局限于基础的副语言感知,缺乏系统的、基于理论的认知框架。我们推出EmoSBench,这是首个基于四分支理论模型构建的SLM情商综合评估基准,涵盖感知、理解、使用和管理情绪四大分支,共包含十个子任务。对EmoSBench的初步评估显示存在巨大差距:即使是GPT-4o-Audio这类领先的专有模型,也仅达到52.6%的准确率,远低于人类基线。为缩小这一差距,我们开发了EmoS,这是一个通过监督微调(SFT)和组相对策略优化(GRPO)优化的专用评估模型。为促进其有效训练,我们构建了EmoDialogue,这是一个双语数据集,通过具有严格定义的情商梯度的响应对提供必要的细粒度监督。同时,我们引入了一种奖励机制,整合了陡峭指数准确率奖励(SEAR)和理由保真度奖励(RFR),以确保精确的序数评分和有效的推理。实验表明,EmoS达到了83.8%的准确率,接近人类水平的表现。此外,对真实、无约束的口语交互的评估验证了其强大的现实泛化能力,为推进情商对话系统奠定了基础框架。
英文摘要
Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a systematic, theory-driven cognitive framework. We introduce EmoSBench, the first comprehensive EI evaluation benchmark for SLMs constructed upon the four-branch theoretical model, covering Perceiving, Understanding, Using, and Managing Emotion across ten sub-tasks. Preliminary assessments on EmoSBench reveal a substantial gap: even leading proprietary models like GPT-4o-Audio achieve only 52.6%, significantly trailing human baselines. To bridge this gap, we develop EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To facilitate its effective training, we curate EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations. Concurrently, we introduce a reward mechanism integrating a Steep Exponential Accuracy Reward (SEAR) and a Rationale Fidelity Reward (RFR) to enforce precise ordinal scoring and valid reasoning. Experiments demonstrate that EmoS reaches 83.8% accuracy, approaching human-level performance. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust real-world generalization, establishing a foundational framework for advancing emotionally intelligent dialogue systems.
CommentsAccepted at ACM Multimedia 2026 (MM '26)