arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经音频编解码器重合成的几何迭代检索

Geometric Iterative Retrieval for Neural Audio Codec Resynthesis

Leo Schmidt-Traub, Frédéric Berdoz, Luca A. Lanzendörfer, Roger Wattenhofer

arXiv 2608.19141首次发表:更新:

发表机构

ETH-DISCO(ETH-DISCO)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对神经音频编解码器从粗令牌重合成高质量音频的问题,提出几何迭代检索方法,在语音和音乐编解码器恢复任务上优于单次令牌预测和单步回归基线。

AI 中文摘要

基于残差向量量化(RVQ)的神经音频编解码器已成为基于令牌的通用音频生成的主要离散表示,但从粗编解码器令牌重合成高质量音频仍是一个未解决的问题,且限制了所有生成此类音频系统的保真度。先前研究将重合成问题视为离散令牌预测与连续回归之间的选择,我们认为这种二分法不完整,并提出几何迭代检索范式,该范式利用RVQ层层次结构本身作为连续码本空间中的自然迭代分解。我们的方法不针对离散词汇表进行分类或回归到单个目标向量,而是在码本的几何空间中执行对比检索。我们在语音和音乐的编解码器恢复任务上评估该方法,结果显示其优于单次通过令牌预测和单步回归基线方法。

英文摘要

Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generation, yet resynthesizing high-quality audio from coarse codec tokens remains an open problem and bounds the fidelity of every system that generates them. Prior work has framed resynthesis as a choice between discrete token prediction and continuous regression. We argue that this dichotomy is incomplete and introduce geometric iterative retrieval, a paradigm that uses the RVQ layer hierarchy itself as a natural iterative decomposition in continuous codebook space. Rather than classifying over discrete vocabularies or regressing to a single target vector, our method performs contrastive retrieval in the codebook's geometric space. We evaluate our method on codec restoration tasks across speech and music, and show improvements over both single-pass token prediction and one-step regression baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑