arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23825cs.CLcs.AI

联邦多语言语音-大语言模型:架构与聚合策略基准测试

Federated Multilingual Speech-LLMs: Architecture and Aggregation Strategy Benchmarking

Jordi Luque, Aleix Sant, Fernando López

首次发表
浏览论文内容

中文总结 AI 辅助

针对多语言语音识别,提出联邦学习基准,评估四种语音-大语言模型架构,发现独立调整学习率和三组件适配可获最佳效果,FedProx有效性依赖架构,为隐私敏感分布式部署提供设计指导。

中文摘要 AI 辅助

我们提出了一个针对多语言自动语音识别(ASR)的联邦学习(FL)综合基准,在Multilingual LibriSpeech数据集上评估了四种语音-大语言模型(Speech-LLM)架构。我们比较了在冻结和未冻结编码器配置下的FedAvg和FedProx,证明优化的学习率对性能至关重要。具体而言,分别为语音编码器、连接器和解码器独立调整学习率可产生最低的错误率,且三组件完全适配(编码器和解码器使用LoRA,连接器进行全量训练)能取得最佳的联邦学习结果。我们观察到FedProx的有效性依赖于架构,在多语言预训练架构中提供了显著优势(例如,在保持编码器固定时,EuroLLM优于TinyLlama);这表明大语言模型骨干网络的容量在调节对异构数据分布的鲁棒性方面起着关键作用。这些发现为在隐私敏感、分布式环境中部署多语言语音-大语言模型提供了具体的设计指导。

英文摘要

We present a comprehensive benchmark of Federated Learning (FL) for multilingual Automatic Speech Recognition (ASR), evaluating four Speech-LLM architectures on the Multilingual LibriSpeech dataset. We compare FedAvg and FedProx across frozen and unfrozen encoder configurations, demonstrating that optimized learning rates are critical for performance. Specifically, independently tuning the learning rates for the speech encoder, connector, and decoder yields the lowest error rates, with full three-component adaptation (LoRA for encoder and decoder, full training for the connector) producing the best FL results. We observe that FedProx efficacy is architecture-dependent, providing notable advantages in multilingual pre-trained architectures (e.g., EuroLLM over TinyLlama when keeping the encoder fixed); this indicates that LLM backbone capacity plays a key role in mediating resilience to heterogeneous data distributions. These findings offer concrete design guidance for deploying multilingual Speech-LLMs in privacy-sensitive, distributed environments.

发表机构

  • Telefónica Innovación Digital, Scientific Research(西班牙电信数字创新研究院)
  • Universitat Politècnica de Catalunya(加泰罗尼亚理工大学)
  • Universidad Autónoma de Madrid(马德里自治大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑