发表机构
School of Communications and Information Engineering, Nanjing University of Posts and Telecommunications(南京邮电大学通信与信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对虚假语音检测中多数据集训练无法适应新欺骗类型及设备存储受限的问题,提出FedCFM框架,利用条件流匹配生成器合成未见欺骗类型嵌入,通过生成重放和知识蒸馏实现持续更新,在相同数据集下取得更低EER,展现强跨域泛化能力。
AI 中文摘要
虚假语音检测(FSD)模型的泛化能力对于实际部署至关重要。现有的多数据集联合训练方法依赖于固定的训练集,无法适应新兴的欺骗类型。尽管持续学习已被探索,但许多方法忽视了单个设备上有限的数据存储,从而限制了实际应用性。为解决这一问题,我们提出了FedCFM,一种基于条件流匹配(CFM)的联邦持续域泛化框架,用于在面临多样且不断演变的欺骗攻击的分布式客户端之间进行协作,而无需共享原始语音数据。每个客户端训练一个基于CFM的生成器,以建模特定欺骗类型的嵌入分布,跨客户端的生成器交换使得能够合成未见过的欺骗类型嵌入,通过生成重放和知识蒸馏实现持续分类器更新。在相同的训练数据集下,FedCFM实现了比评估的集中式和联邦域泛化基线更低的等错误率(EER),展示了强大的跨域泛化能力。代码将在该https URL上发布。
英文摘要
The generalization ability of Fake Speech Detection (FSD) models is crucial for real-world deployment. Existing multi-dataset co-training methods rely on fixed training sets and cannot adapt to emerging spoofing types. Although con-tinual learning has been explored, many approaches overlook limited data storage at individual devices, thereby restricting practical applicability. To address this, we propose FedCFM, a Federated continual domain generalization framework via Conditional Flow Matching (CFM) for collaboration without sharing raw speech data across distributed clients facing diverse and evolving spoofing attacks. Each client trains a CFM-based generator to model spoof-type-specific embedding distributions, and cross-client generator exchange enables synthesis of unseen spoof-type embeddings for continual classifier updating through generative replay and knowledge distillation. With the same training datasets, FedCFM achieves lower EER than the eval-uated centralized and federated domain generalization baselines, demonstrating strong cross-domain generalization. Code will be released on https://github.com/jspycpp/FedCFM.
CommentsAccepted to IEEE SLT 2026