语音语言模型与联邦学习结合实现端到端语音识别:英语和意大利语案例研究
SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies
浏览论文内容
中文总结 AI 辅助
研究将联邦学习应用于基于语音语言模型的端到端语音识别,设计通信高效的联邦优化策略,通过英语和意大利语单语ASR任务评估其有效性与稳定性,进行消融研究分析语音编码器架构影响,为多语言场景部署联邦SpeechLLM奠定基础。
中文摘要 AI 辅助
联邦学习(FL)能在分布式数据源上对自动语音识别(ASR)系统进行隐私保护训练,但其在大规模语音语言模型(SpeechLLMs)中的应用尚待探索。本文首次对基于SpeechLLM的端到端ASR系统的联邦训练进行系统研究。设计了通信高效的联邦优化策略,针对SpeechLLM架构挑战。通过对英语和意大利语单语ASR任务的广泛实证评估,证明联邦方法的有效性和稳定性。还进行消融研究,分析不同语音编码器架构对联邦框架下单语英语ASR性能的影响。结果在降低通信成本的同时实现有竞争力的字错误率,为在实际多语言场景中部署联邦SpeechLLM奠定了实践基础。
英文摘要
Federated learning (FL) enables privacy-preserving training of automatic speech recognition (ASR) systems across distributed data sources, yet its application to large-scale speech language models (SpeechLLMs) remains unexplored. This paper presents the first systematic study of federated training for SpeechLLM-based end-to-end ASR systems. We design a communication-efficient federated optimization strategy tailored to the unique challenges of SpeechLLM architectures, addressing high-dimensional parameter spaces, gradient communication overhead, and computational constraints in distributed settings. Through extensive empirical evaluation on monolingual ASR tasks in English and Italian, we demonstrate the effectiveness and stability of our federated approach compared to centralized training baselines across diverse acoustic conditions and speaking styles. Additionally, we conduct a comprehensive ablation study analyzing the impact of different speech encoder architectures on monolingual English ASR performance within the federated framework, providing insights into optimal model configurations for decentralized training. Our results achieve competitive word error rates while reducing communication costs, establishing practical foundations for federated SpeechLLM deployment in real-world multilingual scenarios.