arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24163eess.AScs.AIcs.LG

非言语发声合成的偏好优化

Preference Optimization for Non-Verbal Vocalization Synthesis

Haoyang Li, Chenglin Xu, Junchuan Zhao, Yuang Cao, Liumeng Xue, Yiwen Guo, Eng Siong Chng

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对支持非言语发声的文本转语音系统,研究偏好优化相关设计,提出NV-CER指标,在Emilia-NV等数据集上验证标准DPO设置的有效性,为相关后训练提供实用见解。

中文摘要 AI 辅助

笑声、咳嗽、叹息等非言语发声(NV)对于富有表现力的文本转语音(TTS)至关重要,但偏好优化在NV生成中的有效性仍未得到充分理解。我们针对支持NV的TTS系统的偏好优化展开系统性研究,重点关注偏好信号、偏好对构建以及基于DPO的优化目标。我们将NV标签视为独立输出符号,计算涵盖言语与非言语内容的加权拼音基础字符错误率,从而提出NV感知字符错误率(NV-CER),该方法无需修改底层优化算法即可实现NV实现的可控优化。在包含18种NV类型的Emilia-NV与增强版NV-Bench数据集上开展的实验,揭示了不同设计选择对NV实现及词汇保真度的影响,并确立了采用标准DPO的有效设置。客观指标、基于大语言模型(LLM)的以及人工评估为我们的发现提供了一致证据,为富有表现力TTS的NV感知后训练提供了实用见解。

英文摘要

Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive TTS, but the effectiveness of preference optimization for NV generation remains poorly understood. We systematically study preference optimization for NV-capable TTS, focusing on preference signals, preference-pair construction, and DPO-based optimization objectives. We formulate an NV-aware character error rate (NV-CER) by treating NV tags as distinct output symbols and computing a weighted pinyin-based CER over both verbal and non-verbal content, enabling controllable optimization of NV realization without modifying the underlying optimization algorithm. Experiments on Emilia-NV and the augmented NV-Bench covering 18 NV types reveal how different design choices affect NV realization and lexical fidelity, and establish an effective setup using standard DPO. Objective, LLM-based, and human evaluations provide converging evidence for our findings, offering practical insights into NV-aware post-training for expressive TTS.

发表机构

  • Nanyang Technological University(南洋理工大学)
  • LIGHTSPEED
  • National University of Singapore(新加坡国立大学)
  • Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑