发表机构
Institute for Advancing Intelligence (IAI), TCG CREST, Kolkata(加尔各答TCG CREST推进智能研究所(IAI))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多说话人对话中合成语音注入的定位问题,提出无需训练的流水线,利用冻结检测器与迟滞解码器,在ASVspoof 5上实现高时间定位精度,并建立零样本基线。
AI 中文摘要
语音克隆欺诈日益依赖于外科手术式注入:在真实对话中,仅将一两句话替换为合成语音。话语级深度伪造检测器对每个片段输出一个真实/伪造标签,无法报告合成语音所在的位置。我们将此形式化为多说话人对话中的时间深度伪造定位(TDLMC),表明当文件同时包含两类样本时,等错误率和最小检测代价函数(min-DCF)是不适定的,并为此场景提出时间度量。我们的贡献是一个无需训练的五阶段流水线,它包装了一个冻结的二元检测器,并在无需重新训练的情况下增加段级输出,使用双阈值迟滞有限状态机解码器将嘈杂的窗口分数转换为连贯的区间。在从ASVspoof 5构建的180个多说话人对话上,该系统使用强骨干网络达到了时间交并比0.90、时间检测率0.95和MS-DCF 0.26,并且在真实语音上的误报率低于6%,在真实的真实多说话人对话(AMI)上低于2%。在相同的流水线下,经过训练的定位器仅将时间IoU提高了约0.04,这限制了放弃监督的代价。在三个冻结检测器和一个在保留校准分割上选择常数的解码器下进行评估,并通过受控分析将残余误报率归因于骨干域差距而非解码器,这为TDLMC提供了第一个零样本基线和可复用的基准。
英文摘要
Voice-cloning fraud increasingly relies on surgical injection: a genuine conversation in which only one or two sentences are replaced by synthetic speech. Utterance-level deepfake detectors emit a single real/fake label per clip and cannot report where the synthetic speech lies. We formalise this as Temporal Deepfake Localisation in Multi-Speaker Conversations (TDLMC), show that equal error rate and min-DCF are ill-posed once a file contains both classes, and propose temporal metrics for this regime. Our contribution is a training-free five-stage pipeline that wraps a frozen binary detector and adds segment-level output with no retraining, using a two-threshold hysteresis finitestate-machine decoder to turn noisy window scores into coherent intervals. On 180 constructed multi-speaker conversations from ASVspoof 5, the system attains temporal intersection-over-union 0.90, temporal detection rate 0.95, and MS-DCF 0.26 with a strong backbone, and its false-alarm rate on genuine speech is below 6%, falling under 2% on genuine real multi-speaker dialogue (AMI). Under an identical pipeline, a trained localiser improves temporal IoU by only about 0.04, bounding the cost of forgoing supervision. Evaluated across three frozen detectors under one decoder whose constants are selected on a held-out calibration split, and with a controlled analysis attributing the residual false-alarm rate to a backbone domain gap rather than to the decoder, this provides the first zero-shot baseline and a reusable benchmark for TDLMC.