发表机构
Dolby(杜比实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种可解释的目的地感知调制与波形恢复流程,先识别调制目的地,再恢复LFO形状和振荡器波形,通过可微分合成器训练,实验表明正确预测目的地至关重要,且Gammatone损失和CQT判别器优于传统MSS损失。
AI 中文摘要
合成器编程是一项具有挑战性的任务,尤其是在配置调制方面:调制是指由低频振荡器(LFO)等调制器对音高、滤波器截止频率和振荡器电平等参数进行时变控制。先前的研究表明,可以通过曲线匹配从干净音频中重建调制,但恢复的参数往往无法迁移到现代合成器上。我们提出了一种可解释的目的地感知调制与波形恢复流程,它模拟了音乐家通常在合成器上重现声音的方式。首先,系统识别被调制的目的地;然后,它恢复与每个目的地相关联的LFO形状,并估计振荡器波形以匹配源音色。训练由一个可微分的合成器支持,该合成器包含噪声振荡器,并且比先前的工作提供更多的调制选项。我们精心训练模型以获得最佳的感知质量,对感知损失和对抗训练进行了广泛的探索。通过客观和主观评估,我们表明正确预测目的地至关重要,我们的模型在面向调制的逆合成任务中表现出色,并且我们的Gammatone损失和CQT判别器在改善感知质量方面显著优于传统使用的MSS损失。我们提供了主观测试中的音频样本。
英文摘要
Synthesizer programming is a challenging task, particularly in configuring modulation: the time-varying control of parameters such as pitch, filter cutoff, and oscillator level by modulators such as Low-Frequency Oscillators (LFOs). Prior work has shown that modulation can be reconstructed from the clean audio through curve matching, yet the recovered parameters are often untransferable to modern synthesizers. We present an interpretable destination-aware modulation and waveform recovery pipeline that mirrors how musicians often recreate sounds on a synthesizer. First, the system identifies the modulated destinations; then it recovers the LFO shape associated with each destination, and estimates the oscillator waveform to match the source timbre. The training is supported by a differentiable synthesizer that incorporates a noise oscillator, and more modulation options than previous work. We carefully train the models for optimal perceptual quality, with extensive exploration of perceptual losses and adversarial training. Through objective and subjective evaluation, we show that predicting the destinations correctly is essential, our model excels in the modulation focused inverse synthesis tasks, and our Gammatone loss and CQT discriminator significantly outperform the traditionally-used MSS loss in improving perceptual quality. We provide audio samples from the subjective test.
CommentsSubmitted to ICASSP2027