arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07549cs.CL

Qwen-Audio-3.0-ASR 技术报告

Qwen-Audio-3.0-ASR Technical Report

Chuanmeng Bian, Daren Chen, Peixin Chen, Zhigao Chen, Zhiyun Fan, Zhifu Gao, Bo Gong, Qing Gu, Jiajun He, Yawei Hu, Yunjie Ji, Jingbei Li, Xiangang Li, Xu Li, Z… 展开作者

Chuanmeng Bian, Daren Chen, Peixin Chen, Zhigao Chen, Zhiyun Fan, Zhifu Gao, Bo Gong, Qing Gu, Jiajun He, Yawei Hu, Yunjie Ji, Jingbei Li, Xiangang Li, Xu Li, Zengxi Li, Zheng Li, Chengdong Liang, Baiji Liu, Ying Liu, Bin Ma, Yiping Peng, Yuezhang Peng, Zhendong Peng, Yu Pu, Yang Shi, Xin Shu, Jian Tang, Biao Tian, Peiyao Wang, Tianzi Wang, Wen Wang, Wupeng Wang, Cheng Wen, Yuzhong Wu, Zijian Xia, Yunchong Xiao, Nan Yang, Jianwei Yu, Jixing Yu, Binbin Zhang, Lei Zhang, Sitong Zhao, Guangdong Zhou, Yuan Zhou, Jianheng Zhuo

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出基于MoE和LLM的Qwen-Audio-3.0-ASR系统,通过统一指令框架处理多语言方言、实体热词及长音频,在工业测试中达到领先性能。

中文摘要 AI 辅助

近年来,自动语音识别(ASR)在三种互补范式的推动下经历了变革性进展:数据扩展、模型扩展以及与大型语言模型(LLM)的深度融合。然而,弥合学术基准性能与实际生产应用之间的差距仍然是一个持续的挑战,尤其是在处理多样化的地域方言、动态实体和热词、长距离上下文信息以及不流畅的自发语音方面。在本报告中,我们介绍了 Qwen-Audio-3.0-ASR,这是一个基于混合专家(MoE)的 LLM 驱动的 ASR 系统,旨在通过统一的指令遵循框架来满足这些生产需求。该模型基于 Qwen 骨干网络构建,并在数千万小时的规模语音数据上进行了训练。Qwen-Audio-3.0-ASR 支持 30 种语言和覆盖八大主要方言区的 16 种中文方言变体的转录。除了多语言和方言识别外,该模型还提供了面向生产的功能,包括行业领域实体识别、分层热词定制、原生单遍转录润色以及长音频上下文建模。我们进一步开发了一个专门的流式变体 Qwen-Audio-3.0-ASR-Streaming,用于对延迟敏感的应用。在中文、英语、多语言和真实世界工业测试集上的广泛评估表明,在广泛的评估条件下,该模型达到了最先进或极具竞争力的识别性能,相对于包括 GPT-4o Transcribe 和 Gemini 3.1 Pro 在内的领先商业和专有系统表现强劲。

英文摘要

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility remains a persistent challenge, particularly in handling diverse regional dialects, dynamic entities and hotwords, long-range contextual information, and disfluent spontaneous speech. In this report, we present Qwen-Audio-3.0-ASR, a Mixture-of-Experts (MoE) LLM-based ASR system designed to address these production demands through a unified, instruction-following framework. The model is built upon the Qwen backbone, and is trained on tens of millions of hours of large-scale speech data. Qwen-Audio-3.0-ASR supports transcription across 30 languages and 16 Chinese dialectal varieties spanning eight major dialect regions. Beyond multilingual and dialectal recognition, the model provides production-oriented capabilities including industry-domain entity recognition, hierarchical hotword customization, native single-pass transcription polishing, and long-audio contextual modeling. We further develop a dedicated streaming variant, Qwen-Audio-3.0-ASR-Streaming, for latency-sensitive applications. Extensive evaluations on Chinese, English, multilingual, and real-world industrial test sets demonstrate state-of-the-art or highly competitive recognition performance across a broad range of evaluation conditions, with strong performance relative to leading commercial and proprietary systems including GPT-4o Transcribe and Gemini 3.1 Pro.

发表机构

  • Alibaba Token Foundry(阿里巴巴代币铸造厂)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑