arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

元学习如何塑造语音深度伪造检测中的LoRA适配器几何结构

How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection

Ivan Kukanov, Janne Laakkonen, Ville Hautamäki

arXiv 2607.22010首次发表:更新:

AI 中文总结

研究语音深度伪造检测中ERM与MLDG的差距,通过固定架构、秩等因素,仅改变目标,用经验Fisher比较二者适配器几何结构,发现目标对适配器投影塑造不同,差距体现在损失相关容量组织方式上,损失感知适配器几何结构可揭示此差异。

AI 中文摘要

用于域泛化的元学习(MLDG)在经验风险最小化(ERM)的基础上改进了分布外语音深度伪造检测,两者都在相同的冻结自监督语音模型上训练低秩适配器。由于架构和适配器容量固定,差距表明训练目标塑造适配器的方式存在差异。我们为此问题引入了一种描述性诊断方法,通过比较ERM和MLDG留下的几何结构来研究。我们用有效秩诊断来表征每个适配器,结果表明目标不会以相同方式重塑所有适配器投影,损失相关更新集中在查询和键投影中,在输出投影中分布更均匀,且在合并更新中也有相同对比,这表明ERM和MLDG之间的差距不仅在于错误率,还在于适配器内损失相关容量的组织方式,而损失感知适配器几何结构是一种观察方式。

英文摘要

Meta-learning for domain generalization (MLDG) improves out-of-distribution speech deepfake detection over empirical risk minimization (ERM) when both objectives train low-rank adapters on the same frozen self-supervised speech model. Because the architecture and adapter capacity are held fixed, this gap points to differences in how the training objective shapes the adapter, yet the field characterizes objectives through error rates rather than through the geometry of the solution they reach. We introduce a descriptive diagnostic for this question: holding architecture, rank, data, and seeds fixed and varying only the objective, we use the empirical Fisher on the finished adapter to compare the geometry that ERM and MLDG leave behind. We characterize each adapter with effective-rank diagnostics that separate where the adapter changes from where those changes matter to the loss, resolved by projection and by depth. Applied to ERM and MLDG, the diagnostic shows that the objective does not reshape all adapter projections alike: the loss-relevant update concentrates in the query and key projections while becoming more distributed in the output projection, consistently across six corpora and most strongly in the upper layers. The same contrast appears in the merged update independently of the low-rank factorization, indicating that it reflects the geometry of the effective update rather than the parameterization. These results show that the gap between ERM and MLDG is not only a difference in error rate, but a difference in how loss-relevant capacity is organized inside the adapter, and that loss-aware adapter geometry is a way to see it.

Comments7 pages, 5 figures, 3 tables. Submitted to SLT 2026 IEEE

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑