发表机构
Indian Institute of Science(印度科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出知晓梵文格律的Vagdhenu TTS系统,采用流匹配TTS主干等组件,解决梵文音系与韵律问题,在比较实验中获近4.6的MOS,已部署并开源相关资源。
AI 中文摘要
我们提出Vagdhenu,这是一个知晓格律(Vrutta)的梵文颂歌(Shloka)转吟唱系统,即一种文本转语音(TTS)系统,可将格律梵文颂歌高保真地转换为其吟唱式的parayana诵念版本。这是一份经验报告,而非新架构的提出。我们采用现成的流匹配TTS主干网络和大规模神经声码器,并添加了忠实梵文吟唱流水线所需的组件:一个前端,通过卡纳达语正字法处理梵文,以避免天城体在印度语言模型中触发的印地语式央音省略;一个遵循梵文微妙音系的前端(包括visarga sandhi及其jihvamuliya和upadhmaniya同位异音、alpaprana和mahaprana的送气对比,以及齿音、卷舌音和腭音咝音的区分);还有一个知晓格律的机制,用于检测格律并依据半参考规则选择完全匹配的参考样本。我们报告了一项塑造了该系统的负面结果:在自填充的流匹配主干网络中,文本侧韵律 conditioner 在架构上是无效的,因为模型会从上下文梅尔频谱中恢复音高,而该嵌入无法获得梯度;参考音频片段和语音引导的再训练是仅有的有效韵律控制手段。我们还报告了跨四个系列(StyleTTS2、VITS2、Matcha-TTS及流匹配主干网络)的比较结果,其中每个较早系列在辅音连缀或韵律方面达到了上限,而五小时时长的克隆模型在专家平均意见得分(MOS)接近4.6的水平上突破了该上限。该系统推出了两个部署版本:包含32章、5183首颂歌的视频语料库(约17.5小时),以及覆盖12部典籍约18000首颂歌的音频应用。我们发布了前端、推理与训练代码、模型权重、单说话人吟唱数据集,以及交互式演示。
英文摘要
We present Vagdhenu, a vrutta (meter) aware shloka-to-chant system for Sanskrit: a text-to-speech system that maps a metrical verse to its chanted parayana recitation at high fidelity. This is an experience report, not a new architecture. We take an off-the-shelf flow-matching TTS backbone and a large-scale neural vocoder, and add the components a faithful Sanskrit chant pipeline needs: a frontend that routes Sanskrit through Kannada orthography to avoid the Hindi-style schwa deletion that Devanagari triggers in Indic models; a frontend that obeys subtle Sanskrit phonology (visarga sandhi with its jihvamuliya and upadhmaniya allophones, the aspiration contrast of alpaprana and mahaprana, and the dental, retroflex, and palatal sibilants kept distinct); and a vrutta-aware mechanism that detects the meter and picks an exactly matched reference under a half-reference rule. We report a negative result that shaped the system: in a self-infilling flow-matching backbone, a text-side prosody conditioner is architecturally inert, because the model recovers pitch from the context mel and the embedding gets no gradient; the reference clip and a voice-steering retrain are the only working prosody levers. We also report a comparative lineage across four families (StyleTTS2, VITS2, Matcha-TTS, and the flow-matching backbone), where each earlier family hit a ceiling on conjuncts or prosody that a five-hour clone cleared at an expert MOS near 4.6. The system shipped two deployments: a 32-chapter, 5183-verse video corpus (about 17.5 hours) and an audio app covering about 18000 verses across 12 books. We release the frontend, inference and training code, weights, a single-speaker chant dataset, and an interactive demo.