arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07423cs.SDcs.LGeess.AS

云端增强型低计算量多通道语音增强

Cloud-Boosted Low-Compute Multi-Channel Speech Enhancement

Xulin Fan, Juan Azcarreta, Ashutosh Pandey, Jesus Alvarez, Ke Tan, Jacob Donley, Ritwik Giri, Buye Xu

首次发表
浏览论文内容

中文总结 AI 辅助

针对可穿戴设备语音增强的计算约束瓶颈,提出融合延迟服务器输出、分层特征增强及协作多通道维纳滤波的云端协作框架,以低额外开销显著提升边缘模型性能。

中文摘要 AI 辅助

低延迟、低计算量的语音增强对具备实时通信需求的可穿戴设备至关重要,但严格的计算约束极大限制了设备端性能。知识增强作为利用更强大的服务器端模型提升边缘模型性能的有效方法已被提出,但其在语音增强领域的性能提升有限。我们提出一种融合三项技术的协作框架:(1)将延迟的服务器输出作为额外输入;(2)分层特征增强,即传递服务器中间表示以指导边缘推理;(3)协作多通道维纳滤波,融合服务器与边缘模型估计的加权协方差矩阵以优化波束成形。实验结果表明,该协作框架在仅增加极少计算开销的情况下,显著优于仅使用边缘模型的基线系统。

英文摘要

Low-latency, low-compute speech enhancement is essential for wearable devices with real-time communication requirements, but strict computational constraints significantly limit on-device performance. Knowledge Boosting has been proposed as an effective approach to improve edge model performance by leveraging a more capable server-side model, but performance gains for speech enhancement have been limited. We propose a collaborative framework incorporating three techniques: (1) delayed server output as additional input, (2) layerwise feature boosting that transfers intermediate server representations to guide edge inference, and (3) collaborative multichannel Wiener filtering, which fuses weighted covariance matrices estimated from both server and edge models for improved beamforming. Experimental results demonstrate that the proposed collaborative framework significantly outperforms the edge-only baseline with minimal additional computational overhead.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Meta Reality Labs Research(Meta现实实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑