arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SpeechCritic:从有限的人类偏好中学习诊断性语音评判器

SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences

Mingyue Huo, Shivam Mehta, Bhavin Jawade, Yinghong Lan, Haoqi Li

arXiv 2609.34582首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Netflix(伊利诺伊大学厄巴纳-香槟分校; Netflix)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SpeechCritic利用约300个人类标注,通过参考条件跨语言设置校准前沿音频模型,实现诊断性语音评判,提升维度级一致性6.3点,并训练7B评判器验证不同训练信号的效果。

AI 中文摘要

人类语音承载着丰富的感觉信息,如情感和说话人身份,然而大多数自动语音质量评判器将其简化为单一的自然度分数。我们研究诊断性语音评判器:给定两个候选,诊断评判器决定哪个更好,以及它们在哪些感知维度(如音色、情感、节奏)上存在差异,并指出哪些可听线索支持其决策。学习此类评判器具有挑战性:专家标注成本高昂,且直接提示前沿音频语言模型生成标签并不可靠:我们的探测揭示了大量错误和不稳定的指令遵循。我们引入SpeechCritic,它在参考条件跨语言设置中仅从约300个人类标注比较中学习诊断评判器。SpeechCritic并非取代前沿模型,而是用这些标签校准它:对每个维度,它选择与人类判断一致的声学测量,将其映射为A/Tie/B概率,并将这些作为非约束性提示连同音频一起传递给模型。与不使用提示的相同模型标注相比,这将维度级人类一致性提高了6.3个百分点,并将与人类Tie率的不匹配降低了10.4个百分点。随后,我们在此监督上训练一个7B评判器,发现不同训练信号塑造不同的评判器行为:SFT确立任务,OPD传递教师的维度级优势和劣势,而RL在人类评分者意见一致的明确比较中帮助最大。值得注意的是,人类听众还发现RL使理由引用更具体、更局部的声学线索,尽管它从未直接奖励理由文本。最后,我们通过在英语-日语和英语-西班牙语上实例化该流程,证明其与语言对无关。综上,这些结果展示了从有限人类偏好到诊断性语音评判器的路径。

英文摘要

Human speech conveys rich perceptual information, such as emotion and speaker identity, yet most automatic speech quality judges reduce it to a single naturalness score. We study diagnostic speech judges: given two candidates, a diagnostic judge decides which is better, along which perceptual dimensions (e.g., timbre, emotion, timing) they differ, and which audible cues support its decision. Learning such judges is challenging: expert annotation is costly, and simply prompting a frontier audio-language model to produce labels is unreliable: our probing reveals substantial errors and unstable instruction following. We introduce SpeechCritic, which learns a diagnostic judge in a reference-conditioned cross-lingual setting from only about 300 human-labeled comparisons. Rather than replacing the frontier model, SpeechCritic calibrates it with these labels: for each dimension, it selects the acoustic measurements that agree with human judgments, maps them to A/Tie/B probabilities, and passes these to the model as non-binding hints alongside the audio. Compared with the same model labeling without hints, this raises dimension-level agreement with humans by 6.3 points and cuts the mismatch with human Tie rates by 10.4 points. We then train a 7B judge on this supervision and find that different training signals shape different judge behaviors: SFT establishes the task, OPD transfers the teacher's dimension-level strengths and weaknesses, and RL helps most on clear-cut comparisons where human raters agree. Notably, human listeners also find that RL makes rationales cite more specific, localized acoustic cues, although it never directly rewards rationale text. Finally, we show that the pipeline is language-pair agnostic by instantiating it on both English-Japanese and English-Spanish. Together, these results demonstrate a path from limited human preferences to a diagnostic speech judge.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑