Anysynth:通过上下文学习和非对称分层引导实现零样本乐器克隆
Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance
浏览论文内容
中文总结 AI 辅助
研究旨在实现零样本乐器克隆,提出基于上下文流匹配的无嵌入神经合成器Anysynth,通过特定条件设定让模型动态检索声学细节,实验显示其性能优越,还提出非对称分层CFG优化可控性,推动了零样本乐器克隆发展。
中文摘要 AI 辅助
零样本乐器克隆旨在仅给定一个简短的[参考音频,参考MIDI]对,就能以未见过的乐器的声学特征生成任意[目标MIDI]序列。现有方法依赖预训练嵌入(如CLAP),将参考音频压缩为固定长度向量,丢弃了忠实音色重建所需的细粒度声学线索。我们提出了Anysynth,一种基于上下文流匹配的无嵌入神经合成器。通过直接在未压缩的参考音频和目标MIDI上对扩散Transformer(DiT)进行条件设定,我们的模型允许自注意力在生成时动态检索声学细节。实验表明,Anysynth在音频质量、音色相似度和旋律贴合度方面优于基于嵌入和自回归的基线。此外,该模型表现出提示长度缩放特性:更长的参考提示会产生更好的音色保真度,这是基于嵌入系统所没有的。为了优化可控性,我们进一步提出了非对称分层CFG,它基于MIDI和参考音色的自然语义-声学依赖在结构上解耦了它们的引导。这种非对称形式避免了梯度冲突,提高了音符准确性和音色保真度,推动了富有表现力的零样本乐器克隆的边界。演示音频可在该https网址获取。
英文摘要
Zero-shot instrument cloning aims to render an arbitrary [Target MIDI] sequence with the acoustic identity of an unseen instrument given only a short [Reference Audio, Reference MIDI] pair. Existing methods rely on pre-trained embeddings (e.g., CLAP) that compress the reference audio into a fixed-length vector, discarding fine-grained acoustic cues essential for faithful timbre reconstruction. We present Anysynth, an embedding-free neural synthesizer based on in-context flow matching. By conditioning a Diffusion Transformer (DiT) directly on the uncompressed reference audio and target MIDI, our model allows self-attention to dynamically retrieve acoustic details at generation time. Experiments show that AnySynth outperforms embedding-based and auto-regressive baselines in audio quality, timbre similarity, and melody adherence. Notably, the model exhibits prompt-length scaling: longer reference prompts yield steadily better timbre fidelity, a property absent in embedding-based systems. To optimize controllability, we further propose Asymmetric Hierarchical CFG, which structurally decouples MIDI and reference-timbre guidance based on their natural semantic-acoustic dependency. This asymmetric formulation avoids gradient conflicts and improves both note accuracy and timbre fidelity, pushing the boundary of expressive, zero-shot instrument cloning. Demo audios are available at https://anysynth-demo.github.io/