arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32016cs.SDcs.AIeess.AS

VoiceNet:超越情感的大规模细粒度语音理解

VoiceNet: Fine-Grained Voice Understanding Beyond Emotion at Scale

Christoph Schuhmann, Robert Kaczmarczyk, Gollam Rabby, Felix Friedrich, Maurice Kraus, Gijs Wijngaard, Kourosh Nadi, Huu Nguyen, Kristian Kersting, Sören Auer

首次发表
浏览论文内容

中文总结 AI 辅助

VoiceNet是首个大规模细粒度语音理解基准,含40种情感和57种说话风格属性标注,并发布VoiceCLAP对比模型,在表示级识别上超越现有CLAP基线,为语音理解提供新工具。

中文摘要 AI 辅助

表现性语音合成已超越表现性语音感知:系统现在能够渲染细粒度的声音表演,而没有任何公开基准能够对其进行评分。针对这一逆问题的大多数基准仅止步于六到九个基本情感类别,且主要基于表演性语音。本文介绍了VoiceNet,一个基于人类标注的、用于在许可协议下的野外语音中进行声音表演理解的表示级基准。VoiceNet包含两个子集:VoiceNet-Emo采用40种情感分类法,每个条目由三位专家评分;VoiceNet-Ext是一个初步子集,对57种说话风格属性进行评分,包括语速、声音紧张度、气声和音域。本文还发布了Emolia,即Emilia语料库的情感标注版本,并包含一个经过精心策划的重平衡子集,该子集通过密集的MOSS-Audio Thinking标注得到丰富。两个语音-文本对比基线在此数据上训练:一个110M参数的VoiceCLAP-Small用于快速大规模数据过滤,以及一个7B参数的VoiceCLAP-Large用于实现最先进的性能。两者均优于现有的CLAP基线,后者在VoiceNet-Emo上接近随机水平。在VoiceNet-Emo上,VoiceCLAP-Large与聚合专家共识的一致性比个体专家之间的一致性更高:这是与多数标签的比较,而非超越人类情感感知的证据。这里评估的所有系统都是语音-文本嵌入模型:VoiceNet对表示级属性识别和检索进行评分,而非端到端的口语对话行为。将未经整理的语音语料库聚类和过滤为涵盖多样说话风格和情感的子集仍然是一个开放的挑战;VoiceCLAP嵌入为此任务提供了一个有前景的工具。VoiceNet、Emolia和VoiceCLAP均可公开用于研究用途。

英文摘要

Expressive speech synthesis has outpaced expressive speech perception: systems now render fine-grained vocal performances that no public benchmark can score. Most benchmarks for this inverse problem stop at six to nine basic emotion categories, largely on acted speech. This paper introduces VoiceNet, a human-annotated representation-level benchmark for voice performance understanding on permissively-licensed in-the-wild speech. VoiceNet has two subsets: VoiceNet-Emo applies a 40-emotion taxonomy with three expert ratings per item, and VoiceNet-Ext, a preliminary subset, scores 57 talking-style attributes including speaking rate, vocal tension, breathiness, and register. The paper also releases Emolia, an emotion-annotated version of the Emilia corpus, with a curated rebalanced subset enriched by dense MOSS-Audio Thinking annotations. Two voice-text contrastive baselines train on this data: a 110M-parameter VoiceCLAP-Small for fast large-scale data filtering and a 7B VoiceCLAP-Large for state-of-the-art performance. Both outperform existing CLAP baselines, which sit near chance on VoiceNet-Emo. On VoiceNet-Emo, VoiceCLAP-Large aligns more closely with the aggregate expert consensus than individual experts agree with one another: a comparison against the majority label rather than evidence of surpassing human emotion perception. All systems evaluated here are voice-text embedding models: VoiceNet scores representation-level attribute recognition and retrieval, not end-to-end spoken-dialogue behaviour. Clustering and filtering uncurated speech corpora into subsets that span diverse talking styles and emotions remains an open challenge; VoiceCLAP embeddings offer a promising tool for this task. VoiceNet, Emolia, and VoiceCLAP are publicly available for research use.

补充信息

↑