arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EmoS:用于评估和对齐口语语言模型中情商的基于理论的框架

EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models

Junyu Wang, Siyuan Zhang, Peiyuan Jiang, Jian Zong, Jingyu Zhang, Tianrui Wang, Yuqin Lin, Zhenghui Chen, Shuqing Xie, Ziyang Ma, Meng Ge, Xiaobao Wang, Longbiao Wang, Jianwu Dang

arXiv 2608.09189首次发表:更新:

发表机构

Tianjin University; Fuzhou University; Shanghai Jiaotong University; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(天津大学; 福州大学; 上海交通大学; 中国科学院深圳先进技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对口语语言模型情商评估缺乏理论框架的问题,构建了EmoSBench基准,开发了基于SFT和GRPO优化的EmoS评估模型,其在EmoSBench上准确率达83.8%,接近人类水平。

AI 中文摘要

尽管遵循指令和听觉理解方面已取得显著进展,但口语语言模型(SLM)的情商(EI)评估仍局限于基础的副语言感知,缺乏系统的、基于理论的认知框架。我们推出EmoSBench,这是首个基于四分支理论模型构建的SLM情商综合评估基准,涵盖感知、理解、使用和管理情绪四大分支,共包含十个子任务。对EmoSBench的初步评估显示存在巨大差距:即使是GPT-4o-Audio这类领先的专有模型,也仅达到52.6%的准确率,远低于人类基线。为缩小这一差距,我们开发了EmoS,这是一个通过监督微调(SFT)和组相对策略优化(GRPO)优化的专用评估模型。为促进其有效训练,我们构建了EmoDialogue,这是一个双语数据集,通过具有严格定义的情商梯度的响应对提供必要的细粒度监督。同时,我们引入了一种奖励机制,整合了陡峭指数准确率奖励(SEAR)和理由保真度奖励(RFR),以确保精确的序数评分和有效的推理。实验表明,EmoS达到了83.8%的准确率,接近人类水平的表现。此外,对真实、无约束的口语交互的评估验证了其强大的现实泛化能力,为推进情商对话系统奠定了基础框架。

英文摘要

Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a systematic, theory-driven cognitive framework. We introduce EmoSBench, the first comprehensive EI evaluation benchmark for SLMs constructed upon the four-branch theoretical model, covering Perceiving, Understanding, Using, and Managing Emotion across ten sub-tasks. Preliminary assessments on EmoSBench reveal a substantial gap: even leading proprietary models like GPT-4o-Audio achieve only 52.6%, significantly trailing human baselines. To bridge this gap, we develop EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To facilitate its effective training, we curate EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations. Concurrently, we introduce a reward mechanism integrating a Steep Exponential Accuracy Reward (SEAR) and a Rationale Fidelity Reward (RFR) to enforce precise ordinal scoring and valid reasoning. Experiments demonstrate that EmoS reaches 83.8% accuracy, approaching human-level performance. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust real-world generalization, establishing a foundational framework for advancing emotionally intelligent dialogue systems.

CommentsAccepted at ACM Multimedia 2026 (MM '26)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑