合成器反演中的离流形鲁棒性:基于联合分布流匹配
Off-manifold robustness in synthesizer inversion with joint distribution flow matching
AI总结:
针对合成器反演中训练数据导致的离流形性能下降问题,提出用多模态连续归一化流建模音频与参数联合分布,利用未配对真实录音训练边缘分布,并在推理时部分加噪,显著提升真实音频反演效果。
AI中文摘要:
近期关于合成器反演的研究表明,生成模型通过显式建模音频到参数映射中的模糊性,其性能优于确定性方法。然而,训练此类模型需要音频-参数对,这些数据通常通过合成器本身渲染采样或预设参数获得。这造成了训练与测试之间的不匹配,可能降低在无真实参数标注的离流形真实世界录音上的性能。为克服这一障碍,我们提出使用多模态连续归一化流对音频和参数的联合分布进行建模,并为每种模态采用独立的噪声调度。该公式允许我们利用配对的合成器数据训练联合密度和条件密度,而无需参数标签的未配对真实录音可单独训练音频边缘分布,使模型暴露于离流形信号。此外,由于模型学习在所有噪声水平下从音频映射到参数,我们发现推理时对音频参考进行部分加噪可改善真实音频的重建,这与降低对分布特定细节的敏感性同时保留粗粒度结构相一致。在Surge XT和Dexed上的评估表明,对联合分布建模显著改善了真实世界和域内音频的反演效果。
英文摘要:
Recent work on synthesizer inversion shows that generative models outperform deterministic approaches by explicitly modeling the ambiguity in mapping audio to parameters. Training such models, however, requires audio-parameter pairs, which are typically obtained by rendering sampled or preset parameters through the synthesizer itself. This creates a train-test mismatch that can degrade performance on off-manifold real-world recordings, for which ground-truth parameter annotations do not exist. To circumvent this obstacle, we propose to model the joint distribution of audio and parameters with a multi-modal continuous normalizing flow using independent noise schedules for each modality. This formulation allows us to train joint and conditional densities with paired synthesizer data, while unpaired real recordings can train the audio marginal alone, exposing the model to off-manifold signals without requiring parameter labels. Further, because the model learns to map from audio to parameters at all noise levels, we find that partially noising the audio reference at inference improves real-audio reconstruction, consistent with reducing sensitivity to distribution-specific detail while preserving coarse structure. Evaluating on Surge XT and Dexed, we find that modelling the joint distribution substantially improves both inversion of real-world and in-domain audio.