arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

追溯起源:神经音频转码中的遗留编解码器识别

Tracing the Origins: Legacy Codec Identification in Neural Audio Transcoding

Wonje Heo, Shinee Youn, Yooshin Kim, Chuck Chae, Donghoon Shin

arXiv 2609.14916首次发表:更新:

发表机构

DGIST(大邱庆北科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对神经音频转码掩盖遗留压缩痕迹的取证空白,提出基于Transformer的框架,利用RVQ序列的层次与时间依赖,实现97%以上准确率的编解码器识别,证明神经转码后传统痕迹仍可追溯。

AI 中文摘要

基于残差向量量化(RVQ)的神经音频编解码器(NAC)通过离散的基于令牌的表示,能够以前所未有的低比特率实现高保真音频分发。然而,这一转变扰乱了传统的取证分析,因为非线性神经转码掩盖了遗留压缩的底层痕迹。本研究界定了这一取证空白,并提出了一种基于Transformer的框架,旨在利用RVQ序列中固有的层次和时间依赖性。通过建模层间因果关系和动态取证重要性,我们的模型有效解耦了从遗留到神经转码过程中叠加的伪影。实验结果显示,在32-128 kbps比特率下,编解码器识别的准确率达到97%以上,并实现了稳健的联合识别性能。这些结果表明,即使在神经转码之后,传统编解码器痕迹仍然存在,支持了神经编解码器感知音频取证的可行性和必要性。

英文摘要

Residual Vector Quantization (RVQ)-based neural audio codecs (NACs) enable high-fidelity audio distribution at unprecedentedly low bitrates through discrete token-based representations. However, this shift disrupts traditional forensics, as non-linear neural transcoding obscures the underlying traces of legacy compression. This study defines the forensic gap and proposes a Transformer-based framework designed to leverage the hierarchical and temporal dependencies inherent in RVQ sequences. By modeling inter-layer causal relationships and dynamic forensic significance, our model effectively disentangles superimposed artifacts from legacy-to-neural transcoding. Experimental results achieve 97%+ accuracy for codec identification and robust joint identification performance across 32-128 kbps. These results demonstrate that traditional codec traces persist even after neural transcoding, supporting the feasibility and necessity of neural-codec-aware audio forensics.

Comments5 pages, 2 figures, to appear Interspeech

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑