AI 中文总结
该研究针对合成器逆问题的两大挑战,提出结合掩码离散扩散与GRPO风格音频域奖励微调的DDSynth-RL方法,在Dexed上验证了其性能优势。
AI 中文摘要
合成器逆问题求解存在两大挑战:1)不同参数配置可产生感知相似的声音;2)参数空间损失函数常无法反映渲染音频的相似性,且合成器作为不可微黑箱,无法应用简单的音频域监督。针对第一个挑战引发的一对多映射问题,我们将合成器逆问题建模为离散合成器参数上的条件生成任务,采用掩码离散扩散作为生成器;该处理方式还避免了自回归模型的固定顺序假设,以及流匹配在建模分类合成器控制时的连续松弛不匹配问题。针对第二个挑战,我们进一步用渲染输出计算的GRPO风格音频域奖励对模型进行微调。在Dexed数据集上的实验表明,经监督训练后,离散扩散模型可与自回归及流匹配基线模型竞争,而基于奖励的微调还进一步提升了域外音频匹配性能。代码与演示可在指定网址获取。
英文摘要
Synthesizer inversion is challenging for two main reasons: 1) Distinct parameter configurations can produce perceptually similar sounds. 2) Parameter-space losses often fail to reflect rendered audio similarity, while the synthesizer being a non-differentiable black box prevents simple audio-domain supervision. To address the one-to-many mapping induced by the first challenge, we formulate synthesizer inversion as conditional generation over discrete synthesizer parameters and use masked discrete diffusion as the generator. This treatment additionally avoids the fixed-order assumption of autoregressive models and the continuous-relaxation mismatch of flow matching when modeling categorical synthesizer controls. To address the second challenge, we further fine-tune the model with GRPO-style audio-domain rewards computed from rendered outputs. Experiments on Dexed show that, after supervised training, the discrete diffusion model is competitive with autoregressive and flow-matching baselines, and reward-based fine-tuning further improves out-of-domain audio matching performance. Code and demos are available at: https://github.com/DDSynth-RL/DDSynthRL.
CommentsAccepted to the 27th International Society for Music Information Retrieval Conference (ISMIR 2026). 8 pages, 3 figures, 1 table