arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VocalRender:面向实际作曲的原生乐谱歌声合成

VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition

Yukun Chen, Tianrui Wang, Zhaoxi Mu, Xinyu Yang, EngSiong Chng

arXiv 2607.27768首次发表:更新:

AI 中文总结

VocalRender是一种原生乐谱歌声合成系统,通过歌词、音高等乐谱信息直接合成歌声,采用交错歌词-音符表示与自回归扩散模型,在2300小时数据集训练后,自然度CMOS较最强基线高0.42分,适配实际作曲工作流。

AI 中文摘要

现有歌声合成系统通常需要预定义时长、显式时长预测或时间对齐的声学引导,这限制了它们与实际作曲工作流的兼容性。我们提出VocalRender,一种原生乐谱系统,可直接根据歌词、音高、符号音符时值和速度合成歌声。它采用交错的歌词-音符表示和自回归扩散模型生成连续声学潜变量,同时预测输出长度,无需显式时长预测。在2300小时的歌声数据集上训练后,VocalRender在域内和域外基准上均实现了强可懂度、强旋律控制和高说话人相似度。值得注意的是,它在自然度CMOS指标上比最强基线高出0.42分,证明了我们提出的原生乐谱架构的有效性。

英文摘要

Existing singing voice synthesis systems often require predefined durations, explicit duration prediction, or time-aligned acoustic guidance, which limits their compatibility with practical composition workflows. We propose VocalRender, a score-native system that directly synthesizes singing from lyrics, pitches, symbolic note values, and tempo. It uses an interleaved lyric--note representation and an autoregressive diffusion model to generate continuous acoustic latents while predicting the output length, eliminating the need for explicit duration prediction. Trained on a 2,300-hour singing dataset, VocalRender achieves strong intelligibility, strong melody control, and high speaker similarity across both in-domain and out-of-domain benchmarks. Notably, it outperforms the strongest baseline by $0.42$ points in naturalness CMOS, demonstrating the effectiveness of our proposed score-native architecture.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑