arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19843cs.SDcs.MM

傅里叶是前沿:用于高保真音乐重建的频率感知自编码

Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction

Kangdi Wang, Yusheng Dai, Jin Xu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对音频自编码器高压缩率下的失效问题,提出ear-VAE2复频谱自编码器及双工感知修正器,在音乐重建任务的多项指标上取得最优表现。

中文摘要 AI 辅助

连续隐变量音频自编码器是隐变量音乐生成器的核心,但高压缩率下的解码器通常会出现三种失效模式:高频损失、相位不一致和立体声像崩塌。这些问题的结构根源在于波形自编码器缺乏显式频率轴,无法进行针对性的逐频带校正。在五个预算匹配的表示方法中,复短时傅里叶变换(complex STFT)实现了最低的全频带和高频频谱距离,可直接获取每个频点的幅度和相位。基于此,我们提出ear-VAE2,一种带有跨通道交互的复频谱自编码器。Spec-SnakeBeta为每个频点学习一个周期性激活函数,采用频率相关初始化,性能优于其他激活变体,且参数数量比完全独立的变体更少。双工感知修正器(Duplex-Aware Refiner)遵循声音定位的双工理论,对幅度和相位进行频带特定校正。在包含546条音轨的Song Describer数据集上,ear-VAE2在7项重建指标中的5项上取得了最佳点估计。双工感知修正器将梅尔距离降低了19.4%,与无约束修正器相比,使用的残差输出维度减少了约45%,同时还降低了频谱距离、空间线索误差,并获得了专业工程师的更高评价。使用ear-VAE2隐变量的下游生成器在全部12项自动指标上取得了更好的点估计。

英文摘要

Continuous-latent audio autoencoders form the backbone of latent music generators, yet decoders at high compression rates commonly exhibit three failure modes: high-frequency loss, phase incoherence, and stereo-image collapse. These share a structural root: waveform autoencoders lack an explicit frequency axis, leaving no handle for targeted per-band correction. Among five matched-budget representations, the complex STFT achieves the lowest full-band and high-frequency spectral distances, providing direct access to magnitude and phase at every bin. Building on this, we present ear-VAE2, a complex-spectral autoencoder with cross-channel interaction. Spec-SnakeBeta learns a periodic activation per frequency bin with frequency-dependent initialization, outperforming other activation variants while using fewer parameters than the fully independent variant. Duplex-Aware Refiner applies band-specific corrections to magnitude and phase following duplex theory of sound localization. On the 546-track Song Describer Dataset, ear-VAE2 achieves the best point estimates on five of seven reconstruction metrics. The Duplex-Aware Refiner reduces Mel Distance by 19.4% and uses ~45% fewer residual-output dimensions than the Unconstrained Refiner, while also lowering spectral distances, spatial-cue errors, and receiving higher ratings from professional engineers. The downstream generator using ear-VAE2 latents achieves better point estimates on all 12 automatic metrics.Demo page is available at https://eps-acoustic-revolution-lab.github.io/EAR_VAE2/.

发表机构

  • Qwen Team, Alibaba(通义千问团队,阿里巴巴)
  • Monash University(莫纳什大学)

机构由 AI 辅助整理,请以论文原文为准。

相关深度报道

↑