半监督联邦语音识别的实用配方:在线伪标签与服务器更新稳定化
A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization
- Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对半监督联邦ASR,提出在线伪标签与服务器更新稳定化的双轴设计,显著缩小与全监督FL的差距。
AI中文摘要:
半监督联邦学习(SSFL)利用教师模型在客户端的未标注数据上生成伪标签,并在服务器上使用少量有标注的种子数据集来训练模型。自动语音识别(ASR)在此场景中尤为脆弱:伪标签错误在输出序列和训练轮次中累积,导致发散,与全监督联邦学习存在较大差距。我们表明,缩小这一差距取决于两个相互耦合的设计轴——教师模型(生成伪标签的模型)和锚点(服务器端在有标注数据上的更新,用于稳定训练)。在教师轴线上,每客户端的在线教师(每个客户端自身演化的模型)会自行发散,但一旦稳定,其性能可与广播全局教师(一个服务器模型,在轮次内固定)相当或更优——在域内表现决定性优势,在域偏移下也具竞争力。随着种子数据增强和在线教师优势缩小,过渡教师(在第r轮从全局切换到在线)可匹配或超越两者。在锚点轴线上,服务器必须在轮次之间持续在有标注数据上训练——否则在线教师会漂移——这种交错训练比种子模型更能决定收敛性。两个轴线不可分割:激进的教师选择仅在锚点稳定训练后才有效,而锚点对数据增强和批大小高度敏感——这些设置决定了服务器注入的输入和梯度噪声量。所需的稳定化程度取决于域,由种子数据的分散性及其与客户端数据的重叠决定。这些发现为ASR训练中的SSFL提供了指导原则,在11对比较中优于最强先前方法9对,域内平均提升20.8%,跨域提升10.0%,缩小了与全监督联邦学习的差距。
英文摘要:
Semi-supervised federated learning (SSFL) trains models on clients' unlabeled data using a teacher to generate pseudo-labels, with a small labeled seed dataset on the server. Automatic Speech Recognition (ASR) is particularly fragile here: pseudo-label errors compound across the output sequence and across training rounds into divergence, leaving a large gap to fully-supervised FL. We show that closing this gap turns on two coupled design axes -- the teacher (which model generates the pseudo-labels) and the anchor (the server-side updates on labeled data that stabilize training). On the teacher axis, a per-client online teacher (each client's own evolving model) diverges on its own, but once stabilized it matches or beats the broadcast global teacher (one server model, fixed within a round) -- decisively in-domain and competitively under domain shift. As the seed grows stronger and the online teacher's advantage narrows, a transitioning teacher (global $\rightarrow$ online at round $r$) matches or beats both. On the anchor axis, the server must keep training on labeled data between rounds -- otherwise the online teacher drifts -- and this interleaving, more than the seed model, governs convergence. The two axes are inseparable: aggressive teacher choices pay off only once the anchor stabilizes training, which is highly sensitive to data augmentation and batch size -- the settings that govern how much input and gradient noise the server injects. How much stabilization is needed is domain-dependent, governed by the dispersion of the seed data and its overlap with client data. These findings yield guidelines for SSFL in ASR training, improving over the strongest prior method on 9 of 11 pairs, by $20.8\%$ on average in-domain and $10.0\%$ cross-domain, narrowing the gap to fully-supervised FL.