arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07981cs.SDcs.LG

干净准确率并不保证溯源鲁棒性:音频归因的前瞻性编解码器压力评估

Clean Accuracy Does Not Guarantee Provenance Robustness: A Prospective Codec-Stress Evaluation of Audio Attribution

Gang Shi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过前瞻性注册实验评估编解码器转码对音频溯源归因的影响,发现干净准确率无法保证部署鲁棒性,编解码器传输可导致显著性能下降,且下降程度依赖条件与表示。

中文摘要 AI 辅助

音频溯源归因——即判断某合成语音由哪个系统生成——在干净基准上报告了接近上限的准确率,然而到达分析人员手中的音频通常已经过转码。我们报告了一项前瞻性注册的测量,该测量在单阶段编解码器传输后进行闭集归因,分析区域在训练任何归因模型之前根据保真度元数据固定。在两个语料库上,在同时的组件级频带下,WavLM-Base+ 的类内损失达到 53.5 [43.5, 63.6] 和 70.3 [63.0, 77.5] 个 Macro-F1 点,W2V2-BERT 2.0 的类内损失达到 61.0 [56.8, 65.1] 和 49.8 [41.6, 57.9] 个 Macro-F1 点。性能下降强烈依赖于条件和表示:在一个类内网格中,WavLM 的损失从 -0.4 到 +53.5 点不等,并且两个编码器在十二个条件中的六个条件下差异超出了预设的 +/-5 点边际。一个干净合格的 ECAPA-TDNN 和一个 Proxy-Anchor 头的性能下降相当,因此该效应并不仅限于某一表示族或弱线性头。注册的匹配保真度比较在此网格上无法估计,并且波形和感知度量对条件的排序不同:MP3 在 8 kbit/s 时在 SI-SDR 上排名中游,但在 PESQ-WB 上排名最后,同时造成最大的损失。对于所测试的任务、语料库、表示和编解码器网格,干净的准确率数字本身并不能表征部署鲁棒性。

英文摘要

Audio provenance attribution - which system produced a synthetic utterance - is reported at near-ceiling accuracy on clean benchmarks, yet audio reaching an analyst has usually been transcoded. We report a prospectively registered measurement of closed-set attribution after single-stage codec transport, with the analysis region fixed from fidelity metadata before any attribution model was trained. On two corpora, in-support losses reach 53.5 [43.5, 63.6] and 70.3 [63.0, 77.5] Macro-F1 points for WavLM-Base+, and 61.0 [56.8, 65.1] and 49.8 [41.6, 57.9] for W2V2-BERT 2.0, under simultaneous component-level bands. Degradation is strongly condition- and representation-dependent: within one in-support grid WavLM losses run from -0.4 to +53.5 points, and the two encoders differ beyond a prespecified +/-5-point margin at six of twelve conditions. A clean-qualified ECAPA-TDNN and a Proxy-Anchor head degrade comparably, so the effect is not confined to one representation family or a weak linear head. The registered matched-fidelity comparison was not estimable on this grid, and waveform and perceptual measures order the conditions differently: MP3 at 8 kbit/s ranks mid-grid on SI-SDR but last on PESQ-WB while causing the largest loss. For the tested tasks, corpora, representations and codec grid, a clean accuracy figure does not by itself characterise deployment robustness.

补充信息

↑