arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25822cs.SDeess.AS

部分音频深度伪造定位的边界与段内学习

Boundary and Intra-Segment Learning for Partial Audio Deepfake Localization

Zhe Ye, Xiangui Kang, Minhua Huang, Kai Wu, Kong Aik Lee, Chng Eng Siong

首次发表
浏览论文内容

中文总结 AI 辅助

针对部分音频深度伪造定位难题,提出边界与段内学习(BISL),联合帧、边界和片段信息建模真实性转换,在PartialSpoof上取得2.52% EER和97.40% F1,优于现有方法。

中文摘要 AI 辅助

部分音频深度伪造仅操纵选定的语音区域,使其难以被定位。现有方法利用边界线索进行部分深度伪造定位,但主要侧重于识别边界位置,而非建模表征真实性转换的特征变化。同时,连续真实片段和伪造片段的内部特征仍未得到充分探索。在本文中,我们提出了边界与段内学习(BISL),该方法引入边界学习来建模相邻帧之间的特征差异,并将真实性转换与一般声学变化区分开来。此外,段内学习捕获连续真实片段和伪造片段的整体特征,同时增强每个片段内的特征一致性。通过联合学习帧、边界和片段信息,BISL能够实现更有效的细粒度部分音频深度伪造定位。在多个定位基准上的实验表明,BISL在PartialSpoof上实现了2.52%的等错误率(EER)和97.40%的F1分数,优于对比方法,同时在HAD上保持了有竞争力的性能,并在LPS上提升了跨数据集性能。代码将在论文被接收后公开提供。

英文摘要

Partial audio deepfakes manipulate only selected speech regions, making them difficult to be localized. Existing methods exploit boundary cues for partial deepfake localization, but primarily focus on identifying boundary positions rather than modeling the feature changes that characterize authenticity transitions. Meanwhile, the internal characteristics of continuous bona fide and spoofed segments remain underexplored. In this paper, we propose Boundary and Intra-Segment Learning (BISL), which introduces boundary learning to model feature differences between adjacent frames and distinguish authenticity transitions from general acoustic variations. In addition, intra-segment learning captures the overall characteristics of continuous bona fide and spoofed segments while enhancing feature consistency within each segment. By jointly learning frame, boundary, and segment information, BISL enables more effective fine-grained partial audio deepfake localization. Experiments on multiple localization benchmarks show that BISL achieves an EER of 2.52\% and an F1-score of 97.40\% on PartialSpoof, outperforming the compared methods, while maintaining competitive performance on HAD and improved cross-dataset performance on LPS. The code will be made publicly available upon acceptance.

发表机构

  • Sun Yat-sen University(中山大学)
  • China Mobile Internet Corporation(中国移动互联网公司)
  • The Hong Kong Polytechnic University(香港理工大学)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑