Deepfake语音检测的最后一公里:产学研经验报告
The Last Mile of Deepfake Speech Detection: An Industry-Academia Experience Report
浏览论文内容
中文总结 AI 辅助
本文结合与Phonexia的三年合作经验,指出Deepfake语音检测落地的障碍,提出商业可用数据集标准等研究与协调建议,为该领域产学研衔接提供参考。
中文摘要 AI 辅助
合成语音检测基准在部分域内评估中报告了低于1%的错误率,但在遭遇未见过的攻击、信道失配和分布偏移时性能会下降。基于与商业说话人识别厂商Phonexia的三年合作,我们报告了构建和部署检测器时遇到的障碍:许多公开基准未获得商业模型开发许可;真实输入并非4秒的干净片段,而是长时长、经编码降级、有时部分合成的录音;当校准后的系统返回2.5的对数似然比时,无人能向客户说明该值对其决策的意义。我们未提出新模型,而是将这些障碍与具体研究及协调建议关联:制定商业可用数据集的共享标准、现实部署基准、非专家可据此采取行动的分数。这些观察来自单个项目,应在其他场景中进行验证。
英文摘要
Synthetic speech detection benchmarks now report sub-1% error rates on some in-domain evaluations, yet performance degrades under unseen attacks, channel mismatch, and distribution shift. Based on a three-year effort with Phonexia, a commercial speaker-recognition vendor, we report barriers encountered while building and deploying a detector. Many public benchmarks are not licensed for commercial model development. Real inputs are not four-second clean clips but long, codec-degraded, sometimes partially synthetic recordings. And when a calibrated system returns a log-likelihood ratio of 2.5, no one can tell the customer what it means for their decision. Rather than proposing a new model, we connect these barriers to concrete research and coordination proposals: shared standards for commercially usable datasets, realistic deployment benchmarks, and scores that non-experts can act on. These observations come from one project and should be tested in other settings.
发表机构
- Brno University of Technology(布尔诺理工大学)
- Phonexia(丰西亚公司)
机构由 AI 辅助整理,请以论文原文为准。