arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14427cs.SDeess.AS

可微数字信号处理混合模型引导的扩散方法用于谐波声音混合中的合成参数估计

Differentiable Digital Signal Processing Mixture Model-Guided Diffusion for Synthesis Parameter Estimation from Harmonic Sound Mixtures

Kengo Takemoto, Tomohiko Nakamura, Hiroshi Saruwatari

首次发表
浏览论文内容

中文总结 AI 辅助

针对DDSP混合模型估计合成参数时缺乏时间平滑性的问题,提出用DDPM引导反向扩散过程,以重建误差为条件,实现时间合理的参数轨迹估计,并在木管和弦乐合奏实验中验证了改进效果。

中文摘要 AI 辅助

可微数字信号处理(DDSP)自编码器通过三类合成参数(基频、响度和音色特征)重建单声道谐波声音。为了在DDSP方法中处理谐波声音的混合,我们先前提出了一种DDSP混合模型(DDSPMM)。该模型将混合信号表示为预训练DDSP自编码器解码器合成的源信号之和。尽管DDSPMM能够直接从混合信号中估计每个源的合成参数,但它并未显式建模合成参数的时间变化,可能产生过度的时域波动。本文提出了一种方法,通过将去噪扩散概率模型(DDPM)纳入基于DDSPMM的估计中,来估计具有时间合理轨迹的合成参数。DDPM被训练为合成参数的生成模型。在估计过程中,所提方法利用观测混合信号与DDSPMM根据当前估计合成的混合信号之间的重建误差来引导DDPM的反向扩散过程。在木管和弦乐合奏上的实验表明,基于DDPM的正则化通过向估计轨迹施加时间合理性,改善了合成参数估计。

英文摘要

A differentiable digital signal processing (DDSP) autoencoder reconstructs a monophonic harmonic sound through three types of synthesis parameters: fundamental frequency, loudness, and timbre features. To handle mixtures of harmonic sounds within the DDSP approach, we have previously proposed a DDSP mixture model (DDSPMM). It represents a mixture as the sum of source signals synthesized by the decoders of pretrained DDSP autoencoders. Although DDSPMM enables direct estimation of synthesis parameters of each source from mixtures, it does not explicitly model temporal variations in the synthesis parameters and can produce excessive temporal fluctuations. In this paper, we propose a method for estimating synthesis parameters with temporally plausible trajectories by incorporating a denoising diffusion probabilistic model (DDPM) into the DDSPMM-based estimation. The DDPM is trained as a generative model of synthesis parameters. During estimation, the proposed method guides the DDPM reverse diffusion process with the reconstruction error between the observed mixture and the mixture synthesized by DDSPMM from the current estimates. Experiments on woodwind and string instrument ensembles showed that the DDPM-based regularization improves synthesis parameter estimation by imposing temporal plausibility on the estimated trajectories.

发表机构

  • The University of Tokyo(东京大学)
  • National Institute of Advanced Industrial Science and Technology (AIST)(产业技术综合研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑