发表机构
DGIST(大邱庆北科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对神经音频转码掩盖遗留压缩痕迹的取证空白,提出基于Transformer的框架,利用RVQ序列的层次与时间依赖,实现97%以上准确率的编解码器识别,证明神经转码后传统痕迹仍可追溯。
AI 中文摘要
基于残差向量量化(RVQ)的神经音频编解码器(NAC)通过离散的基于令牌的表示,能够以前所未有的低比特率实现高保真音频分发。然而,这一转变扰乱了传统的取证分析,因为非线性神经转码掩盖了遗留压缩的底层痕迹。本研究界定了这一取证空白,并提出了一种基于Transformer的框架,旨在利用RVQ序列中固有的层次和时间依赖性。通过建模层间因果关系和动态取证重要性,我们的模型有效解耦了从遗留到神经转码过程中叠加的伪影。实验结果显示,在32-128 kbps比特率下,编解码器识别的准确率达到97%以上,并实现了稳健的联合识别性能。这些结果表明,即使在神经转码之后,传统编解码器痕迹仍然存在,支持了神经编解码器感知音频取证的可行性和必要性。
英文摘要
Residual Vector Quantization (RVQ)-based neural audio codecs (NACs) enable high-fidelity audio distribution at unprecedentedly low bitrates through discrete token-based representations. However, this shift disrupts traditional forensics, as non-linear neural transcoding obscures the underlying traces of legacy compression. This study defines the forensic gap and proposes a Transformer-based framework designed to leverage the hierarchical and temporal dependencies inherent in RVQ sequences. By modeling inter-layer causal relationships and dynamic forensic significance, our model effectively disentangles superimposed artifacts from legacy-to-neural transcoding. Experimental results achieve 97%+ accuracy for codec identification and robust joint identification performance across 32-128 kbps. These results demonstrate that traditional codec traces persist even after neural transcoding, supporting the feasibility and necessity of neural-codec-aware audio forensics.
Comments5 pages, 2 figures, to appear Interspeech