基于潜在算子优化与扩散模型先验的音乐修复
Music Restoration via Latent Operator Optimization and Diffusion Model Priors
浏览论文内容
中文总结 AI 辅助
该研究提出LOUDAR方法,将未知失真建模为可学习潜在算子,结合预训练音频自编码器与潜在扩散模型先验,在歌声、吉他失真修复任务上表现优于退化输入且具竞争力。
中文摘要 AI 辅助
音乐修复旨在从受未知效应、失真或损坏影响的观测录音中恢复干净信号。现有系统常依赖配对训练数据和特定于失真的监督,这限制了其在正向过程未知时的应用。我们提出LOUDAR(音频修复中未知失真的潜在空间优化,Latent-space Optimization of Unknown Distortion for Audio Restoration),这是一种通用修复方法,在预训练音频自编码器的潜在空间中运行,将未知失真建模为可学习的潜在算子。推理时,LOUDAR交替估计干净潜在变量并更新潜在算子参数;无条件潜在扩散模型为干净音频提供先验,通过引导潜在估计向干净录音流形靠拢来正则化该推理过程。由于退化模型是针对每个输入适配的,该方法广泛适用于各类修复问题。我们在歌声效果去除与修复、吉他失真去除任务上评估LOUDAR,结果显示其持续改善退化输入的质量,且在波形和潜在域上与有监督、无监督基线方法具有竞争力。
英文摘要
Music restoration seeks to recover a clean signal from an observed recording degraded by an unknown effect, distortion, or corruption. Existing systems often rely on paired training data and distortion-specific supervision, which limits their use when the forward process is not known in advance. We propose LOUDAR (Latent-space Optimization of Unknown Distortion for Audio Restoration) a general-purpose restoration method that operates in the latent space of a pretrained audio autoencoder and models the unknown distortion as a learnable latent operator. At inference time, LOUDAR alternates between estimating the clean latent variable and updating the latent operator parameters. An unconditional latent diffusion model provides a prior over clean audio and regularizes this inference by steering the latent estimate toward the manifold of clean recordings. Because the degradation model is adapted per input, the approach is broadly applicable across diverse restoration problems. We evaluate LOUDAR on singing voice effect removal and restoration, as well as guitar distortion removal, and show that it consistently improves over degraded inputs and is competitive with supervised and unsupervised baselines in waveform and latent domains.
发表机构
- Brno University of Technology(布尔诺理工大学)
- Aalto University(阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。