基于可微分心理声学损失的高效音频增强方法
Efficient Audio Enhancement with a Differentiable Psychoacoustic Loss
浏览论文内容
中文总结 AI 辅助
本研究提出AEROMamba_P与AEROMamba_PS两种高效音频增强模型,结合Mamba状态空间模型与PAQM可微分损失,在带宽扩展和压缩音频恢复任务中,实现了感知质量提升与计算资源消耗降低的双重优势。
中文摘要 AI 辅助
音频增强旨在提升音频信号的感知质量。本研究首先针对带宽扩展问题,提出了AEROMamba_P,它是AERO超分辨率架构的高效变体,将注意力层和LSTM层替换为Mamba状态空间模型,并结合了基于感知音频质量度量(PAQM)开发的新型可微分感知损失。训练期间,该架构所需的GPU内存比基线少约2至4倍;推理时,它实现了14倍的加速,且仅使用五分之一的GPU内存。在将钢琴数据集和MUSDB18从11.025 kHz上采样至44.1 kHz时,主观听觉测试显示AEROMamba_P的感知质量得分比AERO高出15%。接下来,为处理有损编码高度压缩的音频信号增强,本研究进一步提出了AEROMamba_PS,它采用相同框架,但将STFT重建损失替换为PAQM损失,专门用于增强32 kbps的MP3编码音频。听觉评估显示,在恢复压缩音频时,AEROMamba_PS的质量评分比AEROMamba_P高出52%。这些结果表明,PAQM驱动的训练结合轻量状态空间建模,在带宽受限和压缩音频场景中均能实现高感知质量与计算效率。
英文摘要
Audio enhancement consists of improving the perceived quality of audio signals. Initially, with the aim of addressing bandwidth extension, this work proposes \(AEROMamba_{P}\), an efficient variant of the AERO super-resolution architecture where attention and LSTM layers are replaced by the Mamba state-space model, and which incorporates a newly developed differentiable perceptual loss derived from the Perceptual Audio Quality Measure (PAQM). During training, the architecture requires approximately 2-4x less GPU memory than the baseline; during inference, it achieves a 14x speedup while using only one-fifth of the GPU memory. When upsampling both a piano dataset and MUSDB18 from 11.025 kHz to 44.1 kHz, subjective listening tests show that \(AEROMamba_{P}\) outperforms AERO by 15% in perceived quality scores. Next, to handle the enhancement of audio signals that have been highly compressed by lossy coding, it is further proposed \(AEROMamba_{PS}\), which applies the same framework but replaces STFT reconstruction losses with the PAQM loss, specifically to enhance MP3 encoded audio at 32 kbps. In listening evaluations, \(AEROMamba_{PS}\) achieves 52% higher quality rating than \(AEROMamba_{P}\) when restoring compressed audio. These results demonstrate that PAQM-driven training coupled with lightweight state-space modeling yields high perceptual quality and computational efficiency in both band-limited and compressed audio scenarios.