AI 中文总结
研究神经音频编解码器中用于鲁棒性的水印,将32位消息嵌入连续潜在表示,采用SEANet风格编码器 - 解码器等技术,表征水印载体移动时的权衡,在48kHz语音上提升了EnCodec - 24k比特准确率并降低了PESQ。
AI 中文摘要
神经音频编解码器对音频水印来说是具有挑战性的变换,因为它们会重新编码、量化和重新合成语音。本文研究用于编解码器鲁棒性的连续潜在空间水印。我们不是仅在波形或频谱图上添加水印,而是将一个32位消息嵌入到类似编解码器的语音自动编码器的连续潜在表示中。该流程使用SEANet风格的编码器 - 解码器、基于Conformer的消息嵌入器、RVQ引导的潜在分解以及在信号处理和神经编解码器变换下训练的潜在域检测器。我们表征了在神经解码之前移动水印载体时出现的权衡。在48kHz语音上,EnCodec感知训练将EnCodec - 24k比特准确率从78.8%提高到95.6%和97.1%,而PESQ从3.727降至3.514和3.427。
英文摘要
Neural audio codecs are challenging transformations for audio watermarking because they re-encode, quantize, and resynthesize speech. This paper investigates continuous latent-space watermarking for codec robustness. Instead of adding a watermark only to the waveform or spectrogram, we embed a 32-bit message into the continuous latent representation of a codec-like speech autoencoder. The pipeline uses a SEANet-style encoder-decoder, a Conformer-based message embedder, RVQ-guided latent decomposition, and a latent-domain detector trained under signal-processing and neural-codec transformations. Rather than proposing a final universal watermarking baseline, we characterize the trade-offs that appear when the watermark carrier is moved before neural decoding. On 48 kHz speech, EnCodec-aware training improves EnCodec-24k bit accuracy from 78.8% to 95.6% and 97.1%, while PESQ decreases from 3.727 to 3.514 and 3.427.
Comments7 pages. Submitted to IEEE Spoken Language Technology Workshop (SLT)