发表机构
Columbia University; Meta Superintelligence Labs(哥伦比亚大学; Meta超级智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出CVAM模型,通过CANDOR语料库的人类标注,经监督微调与组相对策略优化,使其语音美学判断优于Gemini 3.1 Pro等模型,为语音美学的人类对齐提供了原则性框架。
AI 中文摘要
我们提出了对话语音美学模型(Conversational Voice Aesthetic Model,CVAM),这是一种用于描述自然对话语境中真实或合成语音回复的语音美学的大型语言模型。给定语境和回复语音,CVAM会描述表征该语音的显著时刻,并预测涵盖性别、音高、语速、情感和表达的9个分类属性。关键挑战在于情感和表达等感知领域,这些领域本质上是主观的,缺乏确定的真实值。因此,我们从CANDOR语料库中收集了约3000个真实和合成回复,每个回复对应约10个人类标注。CVAM在合成美学描述和标签上进行监督微调,随后通过组相对策略优化(Group Relative Policy Optimization)基于人类判断进行优化。实验表明,CVAM与人类听众的一致性优于Gemini 3.1 Pro和开源语音大型语言模型,且优于单人与其余人的一致性。综上,我们证明了将语音美学建立在人类感知基础上的重要性,并提出了一种用于人类对齐的原则性框架。
英文摘要
We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts. Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical attributes spanning gender, pitch, pacing, emotion, and delivery. The key challenge lies in perceptual fields such as emotion and delivery, which are inherently subjective and lack definitive ground truth. Therefore, we collect ~10 human annotations for each of 3k real and synthetic responses derived from the CANDOR corpus. CVAM is supervised finetuned on synthesized aesthetic descriptions and labels, then optimized with Group Relative Policy Optimization on human judgments. Experiments show that CVAM better agrees with human listeners than Gemini 3.1 Pro and open-source speech LLMs, and outperforms single-human-vs.-rest agreement. Together, we demonstrate the importance of grounding voice aesthetics in human perception and propose a principled framework for human alignment.