发表机构
Lab260; BitmanagerAI; MTUCI(Lab260; BitmanagerAI; 莫斯科通信与信息技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过113次实验和嵌入空间分析,发现语音防欺骗中的辅助目标并未带来实质增益,其看似有效源于空间收缩而非不变性,且无法超越交叉熵。
AI 中文摘要
当生成器、编解码器或信道发生变化时,语音防欺骗对策的性能会下降,一种常见的补救措施是使用辅助目标来塑造嵌入空间;然而,这种塑造是否有效在EER(等错误率)这一纯排序指标中是不可见的。我们在五个语料库上进行了113次运行,将七种此类目标与交叉熵进行了比较,使用AASIST3在三个随机种子下以及四种预训练检测器,并直接测量了24次AASIST3运行的嵌入空间。原始增强位移使余弦一致性看起来有效,但增益是一个更小的空间,而非更稳定的空间:按散布归一化后,没有任何配置能持续优于交叉熵。每个训练得到的空间都由预期的二分类单一决策轴主导,其训练集结构无法迁移,且四次运行崩溃为接近恒定的输出,这种输出受到位移的奖励,而EER报告其准确性较差。没有任何辅助目标能在不同架构、语料库和随机种子下保持对交叉熵的优势。
英文摘要
Speech anti-spoofing countermeasures degrade when the generator, codec or channel changes, and a common remedy is an auxiliary objective that shapes the embedding space; whether it does is invisible to EER, a pure ranking metric. We compare seven such objectives with cross-entropy over 113 runs on five corpora, AASIST3 at three seeds plus four pre-trained detectors, and measure the embedding space of the 24 AASIST3 runs directly. Raw augmentation displacement makes cosine consistency look effective, but the gain is a smaller space, not a more stable one: normalised by the spread, no configuration consistently improves on cross-entropy. Every trained space is dominated by the single decision axis expected for two classes, whose training-set structure does not transfer, and four runs collapse to a near-constant output that displacement rewards and EER reports as poor accuracy. No auxiliary objective keeps an advantage over cross-entropy across architectures, corpora and seeds.
CommentsSubmitted to IEEE ICASSP 2027