arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向直播语音合成的多维韵律评判

Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis

Zifan Guan, Longyu Lu, Junan Zhang, Zhizheng Wu, Meiguang Jin, Junfeng Ma

arXiv 2609.20124首次发表:更新:

发表机构

The Chinese University of Hong Kong, Shenzhen; TaoLive-AIGC Team Taobao & Tmall Group of Alibaba(香港中文大学(深圳); 阿里巴巴淘宝天猫集团淘Live-AIGC团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对直播TTS韵律评估成本高、标准多维评估存在判定耦合的问题,提出低成本成对评估器LPJ及解耦版D-LPJ,在准确率、维度独立性及TTS候选选择上均表现优异,可用于细粒度TTS偏好优化。

AI 中文摘要

评估直播语音合成(TTS)需要衡量情感、语调、能量等细粒度、高表现力的韵律特征,而传统的平均意见得分(MOS)预测器无法捕捉这些特征。尽管Gemini等专有大语言模型(LLM)能够评估这些维度,但对于大规模推理和强化学习反馈而言成本过高。为解决这一问题,我们首先提出Live-ProsodyJudge(LPJ,直播韵律评判器)——一种从Gemini蒸馏到Qwen3-Omni的低成本成对评估器。然而,我们发现标准多维评估存在一个关键缺陷:判定耦合。评判器倾向于偷懒地将所有单个维度得分与整体偏好对齐,将丰富的多维评价准则坍缩为单一偏好比特。为解决该问题,我们进一步提出Decoupled-Live-ProsodyJudge(D-LPJ,解耦式直播韵律评判器)。D-LPJ移除整体判定目标以避免盲目跟随,在监督微调(SFT)期间掩码不确定的成对维度,并提出一种新颖的跨度局部GRPO策略,将归一化优势严格应用于对应的推理跨度。在经过精心筛选的人工标注测试集上评估显示,10样本平衡顺序的LPJ点准确率高于单次Gemini调用,而D-LPJ成功生成独立、解耦的维度判定。此外,在8选1 TTS候选选择竞赛中,LPJ选择的语音在85.29%的高置信度案例中进入人类排名前三,证明了其在细粒度TTS偏好优化中的有效性。

英文摘要

Evaluating live streaming speech synthesis (TTS) requires assessing fine-grained, highly expressive prosody such as emotion, intonation, and energy which traditional MOS predictors fail to capture. While proprietary Large Language Models (LLMs) like Gemini can evaluate these aspects, they are too costly for massive inference and reinforcement learning feedback. To address this, we first introduce Live-ProsodyJudge (LPJ), a cost-effective pairwise evaluator distilled from Gemini into Qwen3-Omni. However, we identify a critical flaw in standard multi-dimensional evaluation: verdict coupling. The judge tends to lazily align all individual dimension scores with its overall preference, collapsing a rich multi-dimensional rubric into a single preference bit. To resolve this, we further propose Decoupled-Live-ProsodyJudge (D-LPJ). D-LPJ eliminates the overall verdict target to prevent blind following, masks uncertain pair-dimensions during Supervised Fine-Tuning(SFT), and introduces a novel span-local GRPO strategy that applies normalized advantages strictly to their corresponding rationale spans. Evaluated on highly curated human-annotated test sets, 10 sample balanced-order LPJ achieves higher point accuracy than a single Gemini call, while D-LPJ successfully produces independent,decoupled dimension judgments. Furthermore, in a Best-of-8 TTS candidate selection tournament, the LPJ-selected utterance falls within the human top-3 in 85.29% of high-confidence cases, demonstrating its efficacy for fine-grained TTS preference optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑