arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10345cs.SDcs.CL

PC-Mix:混合语音和环境声音条件下的部分组件音频欺骗检测

PC-Mix: Partial-Component Audio Spoofing Detection under Mixed Speech and Environmental Sound Conditions

Zhenshan Zhang, Xueping Zhang, Linxi Li, Yechen Wang, Ming Li

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对部分音频欺骗检测,提出PC-Mix数据集,解决现有基准差距。构建含真实与部分欺骗环境声音组件并与语音信号混合的音频,建立评估协议与联合学习框架,实验表明匹配条件下训练检测欺骗更有效。

中文摘要 AI 辅助

近期关于部分音频欺骗的研究主要集中在带有欺骗片段时间定位的录音室语音上。然而,这些研究常常忽略了欺骗和真实片段在语音和环境声音组件中同时共存的现实情况。本文提出了PC-Mix,首个用于部分组件欺骗检测的数据集,其中音频组件之一或两者都可能被部分欺骗。在PC-Mix中,首先构建真实和部分欺骗的环境声音组件,并与现有部分欺骗数据集的语音信号混合,生成其中一个或两个组件都可被局部操纵的音频。此设计解决了现有部分欺骗基准中的两个主要差距:语音部分欺骗场景中缺乏现实环境声音以及环境声音组件缺乏部分欺骗检测。我们进一步建立了标准化评估协议,并设计了联合学习框架以优化语音、环境声音和混合音频的欺骗检测。实验突出了混合条件带来的增加的难度。结果表明,在匹配目标条件下训练比直接转移在语音或环境声音组件上训练的模型更有效。

英文摘要

Recent studies on partial audio spoofing mainly focus on studio-recorded speech with temporal localization of spoofed segments. However, these studies often overlook realistic conditions where spoofed and bonafide segments simultaneously coexist across speech and environmental sound components. In this paper, we present PC-Mix, the first dataset for partial-component spoofing detection, where either or both audio components may be partially spoofed. In PC-Mix, bonafide and partially spoofed environmental-sound components are first constructed and mixed with speech signals from an existing partial-spoof dataset, producing audio in which either or both components may be locally manipulated. This design addresses two major gaps in existing partial spoofing benchmarks: the lack of realistic environmental sounds in speech partial spoofing scenarios and the absence of partial spoofing detection for environmental sound components. We further establish standardized evaluation protocols and design a joint learning framework to optimize spoofing detection across speech, environmental sound, and mixed audio. Experiments highlight the increased difficulty introduced by mixed conditions. The results demonstrate that training under matched target conditions is more effective than directly transferring models trained on speech or environmental sound components.

↑