arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向弹性语音推理:基于预训练ASR的无训练唤醒词检测

Toward Elastic Speech Inference: Training-Free Wake-Word Detection from Pretrained ASR

Hwayeon Kim, Youngwon Choi, Hyeonyu Kim

arXiv 2610.01182首次发表:更新:

发表机构

MAUM AI Inc.(MAUM AI 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探索利用预训练ASR模型,通过SliceGPT剪枝提取紧凑编码器,实现无需微调的唤醒词检测,实验显示通道缩减50%时性能稳定。

AI 中文摘要

近年来,ASR的发展日益强调跨不同领域和声学条件的泛化能力。现有方法通常通过额外训练或任务特定模块,将预训练的ASR模型适配到唤醒词(WuW)检测等前端功能上。在本工作中,我们探索了在不进行基于梯度的微调的情况下,使用共享的预训练ASR骨干网络进行唤醒词检测,并考察是否可以利用基于PCA的SliceGPT结构化剪枝方法提取出紧凑的编码器。使用Parakeet-TDT-0.6B-v3和Moonshine-base进行的实验表明,当编码器通道维度缩减50%时,唤醒词检测性能保持相对稳定。这些结果表明,无需微调即可从预训练ASR模型中派生出任务相关的紧凑编码器。

英文摘要

Recent ASR development has placed growing emphasis on generalization across diverse domains and acoustic conditions. Existing approaches typically adapt pretrained ASR models to front-end functions such as wake-up word (WuW) detection through additional training or task-specific modules. In this work, we explore the use of a shared pretrained ASR backbone for WuW detection without gradient-based fine-tuning and examine whether a compact encoder can be extracted using the PCA-based structured pruning approach of SliceGPT. Experiments with Parakeet-TDT-0.6B-v3 and Moonshine-base show that WuW detection performance remains relatively stable when the encoder channel dimension is reduced by 50%. These results suggest that task-relevant compact encoders can be derived from pretrained ASR models without fine-tuning.

CommentsSubmitted to IEEE ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑