AI 中文总结
本研究针对现有神经声码器扩展到空间音频时空间质量下降的问题,提出基于因果GAN的CSAVocoder,通过空间适配器与空间一致性判别器实现联合优化,在大规模数据集上验证其兼具高空间保真度、竞争力音频质量与实时性能。
AI 中文摘要
空间音频声码器可将生成模型产生的梅尔频谱图转换为空间音频波形。大多数神经声码器是为单声道音频设计的,直接扩展到空间音频会因忽略通道间线索而降低空间质量。我们提出CSAVocoder,一种基于因果GAN的空间音频声码器,可联合优化波形保真度和空间渲染。该框架引入了空间适配器,将多通道梅尔频谱图与动态声源-听者姿态信息融合,同时配备空间一致性判别器,用于监督通道间线索。为满足实时要求,我们设计了严格因果的有状态生成器,支持恒定内存开销的高效流式推理。在大规模空间音频数据集上的实验表明,CSAVocoder在具备竞争力的音频质量和实时性能的同时,提升了空间保真度。
英文摘要
Spatial audio vocoders are able to convert mel-spectrograms produced by generative models into spatial audio waveforms. Most neural vocoders are designed for monaural audio, and direct extensions to spatial audio can degrade spatial quality by ignoring inter-channel cues. We present CSAVocoder, a causal GAN-based spatial audio vocoder that jointly optimizes waveform fidelity and spatial rendering. Our framework introduces a Spatial Adaptor that fuses multi-channel mel-spectrograms with dynamic source-listener pose information, together with a spatial consistency discriminator that supervises inter-channel cues. To meet real-time requirements, we design a strictly causal, stateful generator that supports efficient streaming inference with constant memory overhead. Experiments on large-scale spatial audio datasets show that CSAVocoder improves spatial fidelity at competitive audio quality and real-time performance.