arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17931cs.CLcs.MMcs.SD

SpeechSense:面向细粒度语音情感分析的副语言聚焦数据集

SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis

Shicheng Ma, Wenqian Cui, Irwin King

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对现有语音情感分析研究的两大局限,构建了副语言聚焦的细粒度语音情感分析数据集SpeechSense,实验证实具备声学信息的模型性能优于纯文本基线,凸显了声学线索的重要性。

中文摘要 AI 辅助

人工智能的最新进展已彻底改变了语音处理领域,然而有效的语音理解不仅需要识别说话内容,还需洞察说话方式。语音情感分析在解读这类副语言线索方面发挥着关键作用,可应用于招聘、客户服务等多种现实场景。但现有语音情感分析研究存在两大主要局限:其一,主流方法依赖以文本为中心的流程,将自动语音识别与文本分析级联,该过程不可避免地会丢弃韵律、语调等关键声学特征,无法捕捉声学模糊话语中的态度含义;其二,当前基准数据集的标签粒度不匹配,优先关注快乐、悲伤等基本情绪,而非社会敏感性所需的自信、不耐烦等细微人际立场。为解决这些局限,本文提出一种面向细粒度语音情感分析的新型数据集SpeechSense。具体而言,我们定义了一种专门的8类人际立场分类体系,这类立场主要可通过韵律线索(而非仅词汇内容)进行检测,随后基于该分类体系构建了经精心筛选的数据集,构建过程采用高保真语音合成技术并经过严格的人工验证。针对多模态大语言模型(LLMs)、纯文本大语言模型及语音编码器开展的综合实验表明,具备声学信息访问权限的模型始终优于纯文本基线模型。这些结果从经验上验证了声学线索在检测微妙说话者态度方面的首要地位,凸显了SpeechSense的必要性。数据集及补充材料可通过此URL获取。

英文摘要

Recent advances in AI have revolutionized speech processing, yet effective speech understanding requires discerning not just what is said, but how it is said. Speech Sentiment Analysis plays a critical role in decoding these paralinguistic cues for diverse real-world applications such as recruitment and customer service. However, existing Speech Sentiment Analysis research faces two primary limitations. First, dominant approaches rely on text-centric pipelines that cascade Automatic Speech Recognition with text analysis. This process inevitably discards essential acoustic features like prosody and tone, failing to capture attitudinal meanings in acoustically ambiguous utterances. Second, current benchmarks suffer from a mismatch in label granularity, prioritizing basic emotions (e.g., happy, sad) over the nuanced interpersonal stances (e.g., confident, impatient) necessary for social sensitivity. To address these limitations, we propose a novel dataset, SpeechSense, for fine-grained speech sentiment analysis. Specifically, we define a specialized 8-class taxonomy of interpersonal stances detectable primarily through prosodic cues beyond lexical content alone. We then construct a curated dataset based on this taxonomy, built from high-fidelity speech synthesis and rigorous human validation. Comprehensive experiments across multi-modal LLMs, text-only LLMs, and speech encoders demonstrate that models with acoustic access consistently outperform text-only baselines. These results empirically validate the primacy of acoustic cues in detecting subtle speaker attitudes, highlighting the necessity of SpeechSense. Dataset and supplementary materials are available at https://github.com/Sher13cked/SpeechSense.

发表机构

  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑