AI 中文总结
CLASVS是用于歌声合成的连续潜变量自回归模型,通过SCT路由与PSCG方法实现保留旋律的歌词编辑,在普通话基准上优于离散自回归Vevo2,降低46.2%的宏观音素错误率。
AI 中文摘要
参考条件下保留旋律的歌词编辑可在替换歌词的同时保留演唱的节奏、歌手身份与自然度。连续潜变量自回归避免了有限码本,提供带学习型停止机制的分步生成。编辑会产生普通重建不存在的冲突:训练时将参考线索与原始歌词配对,而推理时要求修改后的歌词覆盖与源歌词相关的线索;一个跟随源的补丁可通过自回归历史传播。我们提出CLASVS,其状态-控制-转换(SCT)路由机制保持目标歌词与参考旋律控制的持续性,向因果规划器返回语音进度的语义反馈,并将前一潜变量补丁限制在局部转换内。渐进式状态-控制接地(PSCG)通过无配对编辑、内容一致的普通话重建学习该机制。在两个普通话基准上,CLASVS在所有四项操作上均优于离散自回归Vevo2,且在保持旋律、歌手相似度与感知质量的同时,将宏观音素错误率(macro-PER)降低46.2%。这些结果共同为无乐谱注释的歌词编辑建立了强大的连续自回归操作点,并为更广泛的分步控制奠定基础。音频演示可在我们的项目页面获取:this https URL。
英文摘要
Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoids finite codebooks and offers stepwise generation with learned stopping. Editing creates a conflict absent from ordinary reconstruction: training pairs reference cues with original lyrics, whereas inference asks revised lyrics to override source-lyric-correlated cues; one source-following patch can propagate through AR history. We introduce CLASVS. Its State-Control-Transition (SCT) routing keeps target-lyric and reference-melody controls persistent, returns semantic feedback on phonetic progress to the causal planner, and confines the previous latent patch to the local Transition. Progressive State-Control Grounding (PSCG) learns this contract through paired-edit-free, content-consistent Mandarin reconstruction. On two Mandarin benchmarks, CLASVS improves all four operations over discrete-AR Vevo2 and reduces macro-PER by 46.2%, while maintaining melody, singer similarity, and perceptual quality. Together, these results establish a strong continuous-AR operating point for score-annotation-free lyric edits and a basis for broader stepwise control. Audio demonstrations are available on our project page: https://piedpiperg.github.io/clasvs-demo/.