追踪解码器伪影以实现紧凑合成语音筛查
Tracing Decoder Artifacts for Compact Synthetic Speech Screening
浏览论文内容
中文总结 AI 辅助
针对合成语音检测成本高的问题,提出利用解码器频谱伪影的紧凑梯度提升树前端筛查,以低存储和能耗高效筛选可疑录音,显著降低检测能量。
中文摘要 AI 辅助
近期语音合成与声音克隆的进展加大了对可靠合成语音检测的需求,然而高精度检测器日益依赖大型预训练模型,对每条录音进行调用成本高昂。我们并未替换此类检测器,而是研究一种紧凑的前端筛查方法,该方法以低成本处理所有输入,仅将可疑录音转发至更昂贵的分析环节。为在无需大型学习编码器的情况下实现轻量级筛查,我们利用语音生成操作引入的频谱痕迹。我们分析了学习上采样和逆短时傅里叶变换合成如何产生可预测的频谱伪影,并直接在生成波形中测量其存在。由于这些伪影的强度因生成器而异,我们将解码器引导的频谱测量与短时频谱形状和时间变化的补充描述符相结合,构建一个紧凑的梯度提升树。在七个语音生成器和两个人类语音来源上,所提筛查方法达到了0.021%的等错误率,估计模型存储为151 KiB。当作为模拟级联的第一阶段与一个11.5亿参数检测器配合使用时,它估计将检测能量降低84.4%,同时保持0.050%的合成语音漏检率,展示了解码器引导的声学证据在低成本前端筛查中的潜力。
英文摘要
Recent advances in speech synthesis and voice cloning have increased the need for reliable synthetic-speech detection, yet high-accuracy detectors increasingly rely on large pretrained models that are costly to invoke on every recording. Rather than replacing such detectors, we investigate a compact front-end screen that processes all inputs cheaply and forwards only suspicious recordings for more expensive analysis. To enable lightweight screening without a large learned encoder, we exploit spectral traces introduced by speech-generation operations. We analyze how learned upsampling and inverse short-time Fourier transform synthesis can produce predictable spectral artifacts and measure their presence directly in generated waveforms. Because the strength of these artifacts varies across generators, we combine decoder-guided spectral measurements with complementary descriptors of short-time spectral shape and temporal variation in a compact gradient-boosted tree. Across seven speech generators and two human-speech sources, the proposed screen achieves an equal error rate of 0.021\% with an estimated model storage of 151 KiB. When used as the first stage of a simulated cascade with a 1.15-billion-parameter detector, it reduces estimated detection energy by 84.4\% while operating at a 0.050\% synthetic-speech miss rate, demonstrating the potential of decoder-guided acoustic evidence for low-cost front-end screening.