WST-Graph:用于语音深度伪造检测的保拓扑小波散射前端
WST-Graph: Topology-Preserving Wavelet Scattering Front-End for Speech Deepfake Detection
浏览论文内容
中文总结 AI 辅助
针对语音深度伪造检测,提出WST-Graph前端,通过重构小波散射路径为稀疏调制-载波网格,在保持拓扑的同时减少约60%参数,并在域外基准上取得性能提升。
中文摘要 AI 辅助
声学前端决定了语音深度伪造检测器能够利用的法医线索。小波散射变换(WST)提供具有显式坐标的稳定多尺度系数,但直接展平会模糊路径之间的父子关系。我们引入了WST-Graph,将这些路径重建为稀疏调制-载波网格,用于AASIST图后端。调制级归一化和长度感知的自适应局部注意力池化产生固定的相对时间表示,同时在学习适应之前保留声学轴。这产生了一个具有固定、无参数WST的波形到图接口。我们的配置在保持与AASIST竞争力的同时,使用的可训练参数减少了约60%,并在选定的域外基准上显示出明显增益。这些结果强调了在构建用于基于图的语音深度伪造检测的紧凑、物理接地接口时,保留载波-调制拓扑中父子关系的价值。代码将在该https URL发布。
英文摘要
The acoustic front-end determines which forensic cues a speech deepfake detector can exploit. The wavelet scattering transform (WST) provides stable multiscale coefficients with explicit coordinates, yet direct flattening obscures the parent relation between paths. We introduce WST-Graph, reconstructing these paths as a sparse modulation-carrier grid for an AASIST graph backend. Modulation-level normalization and length-aware adaptive local attention pooling produce fixed relative-time representations while retaining the acoustic axes before learned adaptation. This yields a waveform-to-graph interface with a fixed, parameter-free WST. Our configurations remain competitive with AASIST while using approximately 60% fewer trainable parameters and show clear gains on selected out-of-domain benchmarks. These results underscore the value of preserving parent-child relations within the carrier-modulation topology when constructing a compact, physically grounded interface for graph-based speech deepfake detection. Code will be released at https://github.com/saki-ciallo/wst-graph.
发表机构
- Jinan University(暨南大学)
- College of Cyber Security, Jinan University(暨南大学网络空间安全学院)
机构由 AI 辅助整理,请以论文原文为准。