AI 中文总结
本文提出训练免费的扩散模型声音变形框架smorph,通过解耦时间结构与音色身份,支持三种变形模式,在保持时间结构的同时实现平滑变形,并被音乐家评价为富有表现力和可演奏性。
AI 中文摘要
声音变形,即生成从一个声音身份过渡到另一个声音身份的中间声音,可以成为音乐声音设计的强大工具。现有的基于扩散的变形方法将时间结构与音色身份纠缠在一起,无法在变换其中一个的同时保持另一个固定。我们提出了smorph,一个无需训练的引导框架,它在变换声音是什么的同时保留声音随时间的行为方式,允许用户进行变形,例如在固定音高下从铜管乐器变形到弦乐器。我们在三种变形模式中进行了演示:提示到提示、音频到提示和音频到音频。跨多种数据集的评估表明,smorph能有效产生平滑的变形轨迹,同时相较于基线大幅提高了时间结构和源保留,尽管在某些设置中向目标变换更为保守。在一项探索性案例研究中,音乐家发现smorph轨迹具有表现力和可演奏性,表明结构锚定可以作为乐器互动的有效约束。
英文摘要
Sound morphing, generating intermediate sounds that transition from one sonic identity to another, can be a powerful tool for musical sound design. Existing diffusion-based morphing approaches entangle temporal structure and timbral identity, offering no mechanism to hold one fixed while transforming the other. We present smorph, a training-free guidance framework that preserves how a sound behaves over time while transforming what the sound is, allowing users to morph, for instance from brass to strings at a fixed pitch. We demonstrate across three morphing modes: prompt-to-prompt, audio-to-prompt, and audio-to-audio. Evaluations across diverse datasets show that smorph effectively produces smooth morph trajectories while substantially improving temporal-structure and source preservation over baselines, albeit with more conservative target-ward transformation in some settings. In an exploratory case study, musicians found smorph trajectories to be expressive and playable, suggesting structural anchoring can serve as a productive constraint for instrumental interaction.
CommentsISMIR 2026; Demo page at https://smorphspace.github.io/