没有时间崩溃:解锁冻结音频水印器的鲁棒性与多路复用容量
No Time to Collapse: Unlocking Robustness and Multiplexed Capacity in Frozen Audio Watermarkers
- University of Warwick(华威大学)
- OfSpectrum, Inc.(OfSpectrum 公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对神经音频水印中时间折叠限制鲁棒性与多路复用能力的问题,本文冻结预训练水印器,仅训练低延迟Conformer解码器,在三个水印器上提升攻击恢复与AUROC,并显著改善多路复用联合恢复。
AI中文摘要:
现代神经音频水印系统通常会在时间上重复嵌入消息,然后通过平均、投票或其他固定聚合规则将产生的时序证据折叠成单个载荷。我们认为这种时间折叠限制了鲁棒性和多个载荷的恢复,并且该限制可以在不重新训练底层水印器的情况下解决。我们冻结预训练水印器的编码器和检测器,仅训练一个低延迟的基于Conformer的解码器。该解码器消耗检测器的时序软输出,系统特定的适配器将这些输出池化为一系列窗口级表示,并预测嵌入的消息。在三个冻结的水印器(AURA、AudioSeal和WavMark)上,学习到的解码器提高了对受攻击消息的恢复,并在所有三个水印器上获得了更高的检测AUROC点估计。在受控的全覆盖和部分覆盖多路复用下,在可行片段上经过一到三次链式攻击后,它分别将两个交替载荷词的联合精确恢复提高了9.7-48.0、6.9-17.3和8.2-14.2个百分点。
英文摘要:
Modern neural audio watermarking systems typically embed a message repeatedly across time and then collapse the resulting temporal evidence into a single payload using averaging, voting, or another fixed aggregation rule. We argue that this temporal collapse limits both robustness and the recovery of multiple payloads, and that the limitation can be addressed without retraining the underlying watermarker. We freeze a pretrained watermarker's encoder and detector and train only a low-latency Conformer-based decoder. The decoder consumes the detector's temporal soft outputs, which a system-specific adapter pools into a sequence of window-level representations, and predicts the embedded message. On three frozen watermarkers (AURA, AudioSeal, and WavMark), the learned decoder improves recovery of attacked messages and yields higher detection AUROC point estimates on all three. Under controlled full- and partial-coverage multiplexing, it improves joint-exact recovery of two alternating payload words by 9.7-48.0, 6.9-17.3, and 8.2-14.2 percentage points, respectively, under one to three chained attacks on feasible clips.