将自嵌入音频水印与超低比特率神经编解码器相结合
Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs
- National Institute of Informatics(信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究结合自嵌入音频水印与超低比特率神经编解码器,实现了无训练的音频操纵检测定位与被操纵区域恢复,发现神经编解码器选择是影响性能的主导因素。
AI中文摘要:
语音录音的局部操纵(仅改变话语的局部片段)对内容完整性验证构成重大挑战,因为随着被操纵比例降低,此类编辑的可靠检测与定位会变得更加困难。水印作为一种主动防御方案,可在内容分发前嵌入辅助信息;经典的基于哈希的方案在理想条件下能实现近乎完美的检测与定位,但一旦某段被操纵,原始内容便无法恢复。基于先前的自嵌入音频隐写框架,本研究对理想条件下的主动防御性能展开初步探索,从三个维度拓展研究:帧级定位、多比特最低有效位变体、跨多种超低比特率神经编解码器表示的评估。该框架通过嵌入紧凑的神经编解码器表示而非加密哈希,不仅能恢复被操纵区域,还支持无需欺骗样本的无训练检测与定位。在理想信道条件下针对四种受控操纵类型开展的实验表明,嵌入的有效载荷(以及因此实现的真实内容近似重建)始终能无差错完全恢复。结果还显示,神经编解码器的选择是影响检测与定位性能的主导因素。
英文摘要:
Partial manipulation of speech recordings, where only localized segments of an utterance are altered, poses a significant challenge for content integrity verification, as reliable detection and localization of such edits becomes harder as the manipulated proportion decreases. Watermarking offers a proactive defense alternative by embedding auxiliary information prior to distribution; classical hash-based schemes achieve near-perfect detection and localization under ideal conditions, but the original content cannot be recovered once a segment is manipulated. Building on a prior self-embedding audio steganography framework, this work presents an initial exploration of proactive defense performance under ideal conditions, extending the investigation along three axes: frame-level localization, multi-bit least significant bit variants, and evaluation across multiple ultra-low-bitrate neural codec representations. By embedding a compact neural codec representation rather than a cryptographic hash, the framework additionally enables recovery of the manipulated regions, while supporting training-free detection and localization without spoofed examples. Experiments across four controlled manipulation types under ideal channel conditions show that the embedded payload, and hence an approximate reconstruction of the authentic content, is always fully recovered without bit errors. The results also indicate that the choice of neural codec is the dominant factor for detection and localization performance.