arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10121eess.AScs.SD

DuRe-ST:用于语音深度伪造检测的双关系频谱-时间建模

DuRe-ST: Dual-Relation Spectro-Temporal Modeling for Speech Deepfake Detection

  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Shaole Li, Siqing Qin, Youzhi Tu, Kong Aik Lee

中文总结 AI 辅助

针对语音深度伪造检测忽略频谱-时间协变的问题,提出DuRe-ST模型,联合协方差与图注意力关系建模,在ASVspoof等基准上相对EER降低25.9%-28.4%,且参数增量小。

中文摘要 AI 辅助

以往的语音深度伪造检测器能够通过图注意力自适应地捕获频谱-时间依赖关系,但它们在很大程度上忽略了频谱表示与时间表示之间的协变关系。为解决这一不足,我们从它们的联合协方差构建归一化亲和力图,并应用多项式图滤波来捕获高阶协方差诱导的依赖关系。我们首先开发了Cov-ST,以单独评估基于协方差的关系建模的贡献。尽管它提升了检测性能,但其对多项式阶数的敏感性表明,单独建模协方差关系时鲁棒性有限。因此,我们提出了DuRe-ST,它联合利用协方差诱导关系和图注意力诱导关系,以捕获互补的二阶协变和自适应频谱-时间依赖。实验表明,在ASVspoof基准上,DuRe-ST相比XLSR-AASIST实现了平均相对等错误率(EER)降低25.9%,在四个跨数据集基准上平均降低28.4%,且仅增加了4-8k个可训练后端参数。此外,在四个基准上,其相对EER比现有最强的公开可比模型低2.2-13.2%,同时模型规模小于所考虑的公开模型。

英文摘要

Previous speech deepfake detectors can adaptively capture spectro-temporal dependencies through graph attention, yet they largely overlook the co-variation between spectral and temporal representations. To address this gap, we construct a normalized affinity graph from their joint covariance and apply polynomial graph filtering to capture higher-order covariance-induced dependencies. We first develop Cov-ST to isolate the contribution of covariance-based relational modeling. Although it improves detection performance, its sensitivity to the polynomial order suggests limited robustness when covariance relations are modeled alone. We therefore propose DuRe-ST, which jointly exploits covariance-induced and graph-attention-induced relations to capture complementary second-order co-variation and adaptive spectro-temporal dependencies. Experiments show that DuRe-ST achieves an average relative EER reduction of 25.9% over XLSR-AASIST on the ASVspoof benchmarks and 28.4% across four cross-dataset benchmarks with only 4-8k additional trainable back-end parameters. It further outperforms the strongest publicly available comparison models by 2.2-13.2% in relative EER on four benchmarks, while remaining smaller than the publicly available models considered.

↑