神经变形:神经音频编解码器中序列优化的令牌级变形
Neural Morphing: Sequence-Optimized Token-Level Morphing in Neural Audio Codecs
AI总结:
研究神经音频编解码器中令牌级变形,提出神经变形方法,结合RVQ组转移策略与连续性约束序列匹配器,实现可控音频转换,专注于可部署系统的实现及实时行为。
AI中文摘要:
神经音频编解码器最初是为高保真压缩而开发的;然而,它们的潜在令牌表示和富有表现力的解码器也构成了可控音频转换的强大基础。这项工作引入了神经变形,一种无需训练的令牌域音频效果,它从用户调色板中选择残差向量量化(RVQ)令牌颗粒,并通过预训练的编解码器解码编辑后的流。该方法将分离粗、中、细码本组的RVQ组转移策略与用有界波束搜索取代独立贪婪选择的连续性约束序列匹配器相结合。预期输出是一种可控混合体:源保留节奏组织,而调色板贡献音色和残差细节。我们专注于可部署的VST3/AU系统的实现和实时行为,包括分块渲染、调色板大小缩放和后端健康检查。
英文摘要:
Neural audio codecs were originally developed for high-fidelity compression; however, their latent token representations and expressive decoders also constitute a powerful substrate for controllable audio transformation. This work introduces Neural Morphing, a training-free token-domain audio effect that selects residual-vector-quantized (RVQ) token grains from a user palette and decodes the edited stream through a pretrained codec. The method combines an RVQ-group transfer policy that separates coarse, middle, and fine codebook groups with a continuity-constrained sequence matcher that replaces independent greedy selection with bounded beam search. The intended output is a controlled hybrid: the source preserves rhythmic organization while the palette contributes timbral color and residual detail. We focus on the implementation and realtime behavior of a deployable VST3/AU system, including chunked rendering, palette-size scaling, and backend health checks.