arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11762cs.CL

面向联邦多语言语音-大语言模型的组件感知差分隐私

Component-Aware Differential Privacy for Federated Multilingual Speech-LLMs

  • Telefónica Innovación Digital(西班牙电信数字创新)
  • Universidad Autónoma de Madrid(马德里自治大学)
  • Universitat Politècnica de Catalunya(加泰罗尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Jordi Luque, Fernando López, Aleix Sant

AI总结:

针对语音-LLM联邦学习中跨组件预算崩溃问题,提出α-split双池分配方法,恢复WER效用并以少量噪声开销增强编码器隐私保护。

AI中文摘要:

逐层差分隐私(DP)裁剪通过按参数数量分配每矩阵裁剪预算,提高了联邦学习中的梯度保真度。我们表明,当声学编码器和语言解码器的更新范数相差一个数量级时,这一方法对语音大语言模型(speech-LLMs)失效。单池逐层方法遭受“跨组件预算崩溃”,使词错误率(WER)远偏离全局裁剪水平,或完全导致训练崩溃。当范数不平衡较温和时,自适应单池方法可部分恢复,证实崩溃严重程度与组件间范数比率成正比。我们通过实验在六种逐层方法和三种语音-LLM架构中诊断了根本原因。随后我们提出α-split,一种双池分配方法,将编码器和LLM参数归一化为独立池,并证明联合ℓ2敏感度和原始(ε,δ)-DP保证不变。在架构校准的α下,与平坦DP相比,我们的方法恢复了WER效用,同时以仅+2.6%的LLM噪声开销,为编码器提供了针对基于说话者声音的梯度反转攻击的4.47倍更严格的每组件噪声保护。

英文摘要:

Per-layer differential privacy (DP) clipping improves gradient fidelity in federated learning by allocating per-matrix clipping budgets proportional to parameter count. We show that this recipe breaks for speech large language models (speech-LLMs), when the acoustic encoder and the language decoder differ by an order of magnitude in update norm. Single-pool per-layer methods suffer \emph{cross-component budget collapse}, dragging word error rate (WER) far from flat global clipping or collapsing training entirely. When the norm imbalance is milder, adaptive single-pool methods partially recover, confirming that collapse severity scales with the inter-component norm ratio. We empirically diagnose the root cause across six per-layer methods and three speech-LLM architectures. We then propose \emph{$α$-split}, a two-pool allocation that normalises encoder and LLM parameters into independent pools, and show that joint $\ell_2$ sensitivity and the original $(\varepsilon,δ)$-DP guarantee are unchanged. At architecture-calibrated $α$, our method recovers WER utility compared to flat DP, while granting the encoder $4.47{\times}$ tighter per-component noise protection against speaker voice-based gradient-inversion attacks at only $+2.6\%$ LLM noise overhead.

补充信息

↑