AI 中文总结
该研究针对现有波形扩散模型效率低的问题,提出双路径(DP)架构及DP-DiT、DP-U-Net两个变体,经DCASE和FSD-Kaggle2018数据集实验,3M参数变体性能可媲美50M参数模型。
AI 中文摘要
扩散模型的最新进展已能直接在波形空间生成高保真拟音音效。现有波形扩散模型主要依赖时域架构,如基于CNN的U-Net和DiffWave式模型,或对时间依赖关系建模的频域Transformer。但这些系统通常构建于大模型容量和大量计算成本之上,紧凑高效的波形扩散架构在很大程度上未被探索。本研究提出一种用于波形扩散的双路径(Dual-Path, DP)架构,该架构在时频域的子带和帧轴上沿维度执行自注意力,此DP设计可实现细粒度时频建模同时保持高效率。基于所提出的DP主干,我们开发了两个变体:DP-DiT和DP-U-Net。在DCASE和FSD-Kaggle2018数据集上的实验表明它们具有优异性能,值得注意的是,3M参数的变体实现了可与超过50M参数模型相媲美的性能。音频样本可在该https URL获取。
英文摘要
Recent advances in diffusion models have enabled high-fidelity Foley sound generation directly in the waveform space. Existing waveform diffusion models primarily rely on time-domain architectures, such as CNN-based U-Nets and DiffWave-style models, or frequency-domain Transformers modeling temporal dependencies. However, these systems are typically built with large model capacities and substantial computational costs, leaving compact and efficient waveform diffusion architectures largely underexplored. In this work, we introduce a Dual-Path (DP) architecture for waveform diffusion that performs dimension-wise self-attention along both subband and frame axes in the time-frequency domain. This DP design enables fine-grained temporal-spectral modeling while maintaining high efficiency. Based on the proposed DP backbone, we develop two variants: DP-DiT and DP-U-Net. Experiments on the DCASE and FSD-Kaggle2018 datasets demonstrate their superior performance. Notably, the 3M parameter variant achieves performance comparable to models with more than 50M parameters. Audio samples are available at https://samplesdemo.github.io/DP-Foley/.