发表机构
Besimple AI(贝简单人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出仅用于测试的公开基准VocalAffectBench,评估AI音频模型的语音情感识别能力,发现现有基线模型仅能提取部分情感信号,离散情感识别仍较脆弱。
AI 中文摘要
语音产品日益需要语音中存在但转录文本中缺失的情感线索。我们推出VocalAffectBench,这是一个仅用于测试的公开基准,用于评估AI音频模型能否从原始音频中识别表达的语音情感。该基准包含来自51个说话人账号的273个人类录制的英文WAV片段,总时长1.95小时,涵盖愤怒、厌恶、恐惧、快乐、中性、悲伤、惊讶七个类别,每个类别各39个片段。所有基线模型仅从音频进行评估,不使用转录文本或上下文元数据。在6个已发布的基线模型中,平均准确率为35.5%;最强基线gemini_3_5_flash在七分类任务上达到46.5%,高于14.3%的随机基线,但仍远未实现稳健的情感识别。我们还进行了二级效价分组分析,将标签映射为正、中、负三类,排除了效价模糊的惊讶类别,此较粗粒度视图下的总准确率为50.9%。不同类别间的性能差异显著:按召回率计算,中性的识别最可靠,基线平均达75.6%,而惊讶和恐惧仅分别为10.7%和15.4%。这些结果表明,所评估的基线模型可从语音中提取部分情感信号,但离散的表达情感识别仍较为脆弱,尤其在语音智能体工作流中通常最为重要的非中性情感方面。
英文摘要
Voice products increasingly need affective cues that are present in speech but absent from transcripts. We introduce VocalAffectBench, a public, test-only benchmark for evaluating whether AI audio models can identify expressed vocal emotion from raw audio. The benchmark contains 273 human-recorded English WAV clips from 51 speaker accounts totaling 1.95 hours across seven labels: angry, disgusted, fearful, happy, neutral, sad, and surprised, with 39 clips per class. All baselines are evaluated from audio alone, without transcripts or contextual metadata. Across six released baselines, average accuracy is 35.5%. The strongest baseline, gemini_3_5_flash, reaches 46.5% on the seven-way task, above the 14.3% random baseline but far from robust emotion recognition. A secondary valence-bucket analysis maps labels into positive, neutral, and negative classes, excluding surprised because its valence is ambiguous. Aggregate accuracy under this coarser view is 50.9%. Performance is highly uneven across classes. By recall, neutral is identified most reliably at 75.6% averaged across baselines, while surprised and fearful reach only 10.7% and 15.4%, respectively. These results show that the evaluated baselines can extract some affective signal from speech, but discrete expressed-emotion recognition remains fragile, especially for non-neutral emotions that are often most important in voice agent workflows.
Comments6 pages, 6 tables. Benchmark, baseline predictions, and aggregate results are publicly available