arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GS-Codec:用于神经音频编码的高斯溅射瓶颈

GS-Codec: A Gaussian-Splatting Bottleneck for Neural Audio Coding

Ron Aluf, Alon Canfi, Eliya Nachmani

arXiv 2610.04651首次发表:更新:

发表机构

School of Electrical and Computer Engineering; Ben-Gurion University of the Negev(电气与计算机工程学院; 内盖夫本-古里安大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GS-Codec提出用参数化信号分解替代量化器作为神经音频编码瓶颈,通过高斯溅射拟合一维潜在表示,实现训练后比特率控制,性能媲美EnCodec和DAC。

AI 中文摘要

神经音频编解码器将波形压缩成紧凑的离散令牌,这些令牌支撑着语音语言模型、实时通信和大规模音频存储。几乎所有主流设计,包括残差向量量化、有限标量量化和单码本变体,都遵循VQ-VAE模板,通过学习的码本或固定的标量网格对编码器潜在表示进行划分。我们质疑这种划分是否必要。我们提出了GS-Codec,一种神经语音编解码器,其瓶颈是参数化信号分解而非量化器。我们将高斯溅射从3D场景重建适配到一维潜在表示。一个内部优化循环将每个编码器段拟合为1D高斯原语的加权和。解码器随后从渲染的总和中重建波形。为了避免推理时这种迭代内部循环的成本,我们额外训练了一个轻量级GS预测网络,该网络通过单次前向传播回归原语参数。编码器和解码器通过内部循环进行端到端训练,训练流程中没有任何量化器:瓶颈是分解本身,标量量化仅在训练后应用于拟合参数。我们的表示不依赖离散码本阶段来控制比特率,而是暴露了细粒度的率-质量权衡:单个训练检查点通过改变原语数量和每参数位深度来支持训练后比特率控制,无需重新训练。GS-Codec在可比比特率下,在说话人相似度(SIM)、可懂度(STOI)和感知质量(UTMOS)方面匹配或超过成熟的开源编解码器(如EnCodec和DAC),同时实现相当的语义性能(WER)。代码和音频样本可在该https URL获取。

英文摘要

Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook variants, follows the VQ-VAE template by partitioning the encoder latent through a learned codebook or a fixed scalar grid. We ask whether this partition is necessary. We introduce GS-Codec, a neural speech codec whose bottleneck is a parametric signal decomposition rather than a quantizer. We adapt Gaussian splatting from 3D scene reconstruction to one-dimensional latents. An inner optimization loop fits each encoder segment as a weighted sum of 1D Gaussian primitives. The decoder then reconstructs the waveform from the rendered sum. To avoid the cost of this iterative inner loop at inference time, we additionally train a lightweight GS Predictor Net that regresses the primitive parameters in a single forward pass. The encoder and decoder are trained end-to-end through the inner loop, with no quantizer anywhere in the training pipeline: the bottleneck is the decomposition itself, and scalar quantization is applied only post-training to the fitted parameters. Rather than relying on discrete codebook stages for bitrate control, our representation exposes a fine-grained rate-quality tradeoff: a single trained checkpoint supports post-training bitrate control by varying the number of primitives and the per-parameter bit depth, with no retraining required. GS-Codec matches or exceeds well-established open-source codecs such as EnCodec and DAC on speaker similarity (SIM), intelligibility (STOI), and perceptual quality (UTMOS) at comparable bitrates, while achieving comparable semantic performance (WER). Code and audio samples are available at https://ronaluf.github.io/gs-codec/

CommentsAccepted to Neurips 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑