部分语音欺骗中的时间锚点与编辑敏感性:一项受控研究
Temporal Anchors and Editing Sensitivity in Partial Speech Spoofing: A Controlled Study
浏览论文内容
中文总结 AI 辅助
本研究通过受控实验诊断部分语音欺骗检测中的时间锚点与边界一致性训练,发现其效果不统一,并报告了帧级定位结果与回顾性评估方法。
中文摘要 AI 辅助
部分欺骗检测器必须在接受良性编辑的同时拒绝合成内容。我们使用冻结的WavLM特征、非语言单元控制和源锁定评估来诊断时间锚点和边界一致性训练。真实-真实和真实-伪造拼接对比了编辑误报与合成内容漏检;它们并未隔离出唯一的因果伪影。使用第6层特征时,电话单元在PartialSpoof评估帧EER上高于软帧(11.18%对比9.59%)。在500例回顾性PS评估案例中,增强减少了真实拼接误报;合成核心漏检在5% PS开发FPR下增加,但在1%下未得到显著证明。一项200例的Llama A后续研究保留了误报减少的效果,但未确立漏检率的增加。构造案例排名可以改善,而固定阈值漏检上升。一致性未带来统一增益,编码器层排名在不同语料库间变化。我们报告了使用三种随机种子(包括词单元)的帧和事件定位,并将PS评估视为回顾性。审计代码、选定的训练脚本和结果摘要可在该https URL获取。
英文摘要
Partial-spoof detectors must reject synthetic content while accepting benign edits. We diagnose temporal anchors and boundary-consistency training using frozen WavLM features, nonlinguistic unit controls, and source-locked evaluation. Genuine-genuine and genuine-fake splices contrast editing false alarms with synthetic-content misses; they do not isolate a unique causal artifact. With layer-6 features, phone units have higher PartialSpoof evaluation frame EER than soft frames (11.18% versus 9.59%). In 500 retrospective PS-eval cases, augmentation reduces genuine-splice false alarms; synthetic-core misses increase at 5% PS-development FPR but not demonstrably at 1%. A 200-case Llama A follow-up retains the false-alarm reduction but does not establish a miss-rate increase. Constructed-case ranking can improve while fixed-threshold misses rise. Consistency gives no uniform gain, and encoder-layer rankings change across corpora. We report frame and event localization with three seeds, including word units, and treat PS evaluation as retrospective. Audit code, selected training scripts, and result summaries are available at https://github.com/mysxs/partial-spoof-diagnostics.
发表机构
- Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
- School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
机构由 AI 辅助整理,请以论文原文为准。