多速率带宽扩展:基于神经音频编解码器的令牌补全
Multi-Rate Bandwidth Extension by Token Completion in Neural Audio Codecs
- LTCI, Télécom Paris, Institut Polytechnique de Paris(巴黎电信学院,巴黎综合理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出将带宽扩展视为音频令牌预测问题,通过解纠缠神经音频编解码器与Transformer语言模型联合设计,实现高质量重建,强调表示感知设计的重要性。
AI中文摘要:
带宽扩展,即从音频信号的低通版本重建其高频成分的任务,是音频处理中一个长期存在的问题。在本工作中,我们将带宽扩展框定为音频令牌预测问题,从而扩展了神经架构的最新进展。具体而言,我们在由解纠缠神经音频编解码器产生的离散表示上训练一个基于Transformer的语言模型,其中解纠缠由输入信号的谐波-打击乐分解引导,突出了与带宽扩展特别相关的频谱结构。我们的方法引入了一种新颖的编解码器设计,该设计明确考虑了下游令牌预测任务,使得编解码器结构与Transformer建模之间能够更有效地耦合。这种联合设计能够根据客观指标和主观评估,高质量地重建原始信号。这些结果凸显了将编解码器解纠缠和表示学习与生成建模阶段对齐的重要性,并展示了全局、表示感知设计在推进带宽扩展方面的潜力。
英文摘要:
Bandwidth extension, the task of reconstructing the high-frequency components of an audio signal from its low-passed counterpart, is a long-standing problem in audio processing. In this work, we extend recent advances in neural architectures by framing bandwidth extension as an audio token prediction problem. Specifically, we train a transformer-based language model on the discrete representations produced by a disentangled neural audio codec, where the disentanglement is guided by a Harmonic-Percussive decomposition of the input signals, highlighting spectral structures particularly relevant for bandwidth extension. Our approach introduces a novel codec design that explicitly accounts for the downstream token prediction task, enabling a more effective coupling between codec structure and transformer modeling. This joint design yields high-quality reconstructions of the original signal, as measured by both objective metrics and subjective evaluations. These results highlight the importance of aligning codec disentanglement and representation learning with the generative modeling stage, and demonstrate the potential of global, representation-aware design for advancing bandwidth extension.