发表机构
National Research Council Canada; The Chinese University of Hong Kong; Shanghai Jiao Tong University(加拿大国家研究委员会; 香港中文大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对语音深度伪造检测中单一分数缺乏可解释性的问题,提出携带多线索的可审计决策记录,经后期校准在ASVspoof 5上显著降低等错误率,并保留证据以支持审查。
AI 中文摘要
语音深度伪造能够模仿说话者的声音,足以欺骗听者和自动化系统。这推动了语音深度伪造检测的强劲进展,但大多数检测器仍以每个话语一个分数作为输出。该分数对排名系统有用,但对于边界项为何应被信任、推迟或审查,它提供的信息很少。两个话语可能因不同原因落入相同的分数区间,例如,被动证据与检索证据不一致,或键控探针不可用。我们探讨最终决策能否在保留这一来源信息的同时保持标量形式。我们通过一个可审计的决策记录来回答这个问题,该记录将四个对齐的线索带入后期校准步骤:一个被动检测器分数、一个基于标记导数的条件键控探针分数、检索支持度以及说话者画像边际,同时附带明确的差异坐标。在包含4,080个样本的ASVspoof 5 Track 1匹配子集上,固定的检索增强规则相较于仅检索证据有所改进,等错误率从15.84%降至11.91%,而对完整记录进行后期校准后,等错误率达到8.43%。在33.75%的审查预算下,暴露的线索并集覆盖了校准模型82.85%的错误。最佳被动WavLM运行仍达到6.71%的等错误率,因此我们不将决策记录作为更强的独立检测器呈现。其贡献在于保留每个表面化话语背后的证据,同时仍产生一个用于阈值设定和审查的操作分数。
英文摘要
Speech deepfake detectors usually emit one score per utterance, but a borderline score does not reveal why two examples differ during retrospective error analysis. We ask whether a final score can be calibrated from component fields while keeping those fields visible for inspection. We build a decision record with a passive detector score and a score from a probe applied to a marked copy. It also includes retrieval support held out of the evaluated family, a margin from a support-set profile, and raw neighbor closeness. A cross-fit calibrator combines these fields and two differences between raw scores into one final score. On matched ASVspoof development data, the calibrated record reduces equal error rate (EER) by 3.48 percentage points relative to the fixed retrieval-augmented rule. It reaches 8.43% EER, whereas a passive WavLM baseline reaches 6.71% on the same subset. The record is therefore not the strongest detector in this comparison. Its value is to retain inspectable component fields while producing one scalar score for retrospective diagnosis.