arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02918eess.AS

基于可微分心理声学损失的高效音频增强方法

Efficient Audio Enhancement with a Differentiable Psychoacoustic Loss

Wallace Abreu, Bernardo V. Miranda, Luiz W. P. Biscainho

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出AEROMamba_P与AEROMamba_PS两种高效音频增强模型,结合Mamba状态空间模型与PAQM可微分损失,在带宽扩展和压缩音频恢复任务中,实现了感知质量提升与计算资源消耗降低的双重优势。

中文摘要 AI 辅助

音频增强旨在提升音频信号的感知质量。本研究首先针对带宽扩展问题,提出了AEROMamba_P,它是AERO超分辨率架构的高效变体,将注意力层和LSTM层替换为Mamba状态空间模型,并结合了基于感知音频质量度量(PAQM)开发的新型可微分感知损失。训练期间,该架构所需的GPU内存比基线少约2至4倍;推理时,它实现了14倍的加速,且仅使用五分之一的GPU内存。在将钢琴数据集和MUSDB18从11.025 kHz上采样至44.1 kHz时,主观听觉测试显示AEROMamba_P的感知质量得分比AERO高出15%。接下来,为处理有损编码高度压缩的音频信号增强,本研究进一步提出了AEROMamba_PS,它采用相同框架,但将STFT重建损失替换为PAQM损失,专门用于增强32 kbps的MP3编码音频。听觉评估显示,在恢复压缩音频时,AEROMamba_PS的质量评分比AEROMamba_P高出52%。这些结果表明,PAQM驱动的训练结合轻量状态空间建模,在带宽受限和压缩音频场景中均能实现高感知质量与计算效率。

英文摘要

Audio enhancement consists of improving the perceived quality of audio signals. Initially, with the aim of addressing bandwidth extension, this work proposes \(AEROMamba_{P}\), an efficient variant of the AERO super-resolution architecture where attention and LSTM layers are replaced by the Mamba state-space model, and which incorporates a newly developed differentiable perceptual loss derived from the Perceptual Audio Quality Measure (PAQM). During training, the architecture requires approximately 2-4x less GPU memory than the baseline; during inference, it achieves a 14x speedup while using only one-fifth of the GPU memory. When upsampling both a piano dataset and MUSDB18 from 11.025 kHz to 44.1 kHz, subjective listening tests show that \(AEROMamba_{P}\) outperforms AERO by 15% in perceived quality scores. Next, to handle the enhancement of audio signals that have been highly compressed by lossy coding, it is further proposed \(AEROMamba_{PS}\), which applies the same framework but replaces STFT reconstruction losses with the PAQM loss, specifically to enhance MP3 encoded audio at 32 kbps. In listening evaluations, \(AEROMamba_{PS}\) achieves 52% higher quality rating than \(AEROMamba_{P}\) when restoring compressed audio. These results demonstrate that PAQM-driven training coupled with lightweight state-space modeling yields high perceptual quality and computational efficiency in both band-limited and compressed audio scenarios.

补充信息

↑