发表机构
University of Rochester; Sony Computer Science Laboratories; National Taiwan University; Embertone(罗切斯特大学; 索尼计算机科学实验室; 国立台湾大学; Embertone)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出潜在扩散框架VIOLET,结合DiT与整流流,利用新整理的CSV-TD数据集训练,实现高保真可控小提琴合成,性能优于现有系统且接近顶级商业虚拟乐器。
AI 中文摘要
乐器的神经合成技术有望革新当前使用拼接合成与样本库的实践,但多数研究聚焦于钢琴合成与富有表现力的演奏生成,针对小提琴这类连续关节乐器的研究较少,更未涉及演奏技巧与动态的呈现。本文提出VIOLET,一种用于可控小提琴合成的潜在扩散框架,采用带整流流的扩散Transformer(DiT),可根据MIDI音符、演奏技巧及连续动态生成高保真音频。为训练VIOLET,除使用少量现有数据集外,还整理了名为CSV-TD的新数据集,包含39小时48kHz的合成音频及时间对齐的MIDI音符、音符级技巧、连续动态曲线标注。客观与主观评估显示,VIOLET合成的小提琴演奏具有高技巧贴合度、准确的音高与时序对齐及良好的动态控制,其在技巧清晰度、自然度与动态跟随方面优于当前最先进的神经小提琴合成系统,且接近顶级商业虚拟乐器的表现。
英文摘要
Neural synthesis for musical instruments has the potential to revolutionize current practices that use concatenative synthesis and a sample library. However, most research focused on piano synthesis and expressive performance generation; little work has been done on continuously articulated instruments like the violin, let alone rendering them with playing techniques and dynamics. We present VIOLET, a latent-diffusion framework for controllable violin synthesis, which uses a Diffusion Transformer (DiT) with rectified flow to synthesize high-fidelity audio from MIDI notes, playing techniques, and continuous dynamics. To train VIOLET, in addition to using a few existing datasets, we curate a new dataset named CSV-TD, which contains 39 h of 48 kHz synthetic audio and time-aligned annotations of MIDI notes, note-level techniques, and continuous dynamics curves. Objective and subjective evaluations show that VIOLET synthesizes violin performances with high technique adherence, accurate pitch and timing alignment, and good dynamics control. It outperforms the current state-of-the-art neural violin synthesis system and approaches a top commercial virtual instrument in terms of technique clarity, naturalness, and dynamics following.
CommentsAccepted at ISMIR 2026; Code and Demo available at https://github.com/User-tian/VIOLET