发表机构
Northwestern Polytechnical University; School of Intelligence Science and Technology, Nanjing University; Nanyang Technological University; Beijing Huiting Technology Co., Ltd.(西北工业大学; 南京大学智能科学与技术学院; 南洋理工大学; 北京慧亭科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文综述2026年中国之声挑战赛,该赛设汉语多方言识别与ASR两项任务,用约320小时语音数据,分析28支参赛团队结果,为多方言语音处理系统研发评估提供指导。
AI 中文摘要
本文对2026年中国之声挑战赛进行了综述,该挑战赛旨在为汉语方言语音处理建立统一的任务定义和评估条件,推动多方言识别与自动语音识别(ASR)技术发展。挑战赛涵盖16个方言类别,设置两项任务:汉语多方言识别和汉语多方言自动语音识别(ASR)。挑战赛使用约320小时的语音数据,分为参考集、开放评估集和隐藏评估集。两项任务使用相同的评估音频,各包含受限数据赛道和开放数据赛道。本文描述了任务设置、数据、评估指标及Qwen3-ASR-1.7B基线模型,分析了排行榜结果与提交系统。共有28支团队提交结果,17支提供系统报告,15支团队的系统通过合规审查并纳入分析。多数符合要求的系统性能优于基线模型,两项任务的官方前三名排名在隐藏评估集上保持不变。方言层面结果显示,识别准确率较高的类别通常ASR错误率较低,尽管两项任务评估的是相关但不同的能力。领先的识别系统通常利用方言区分性的声学表征,而领先的ASR系统则强调数据归一化、数据增强和辅助连接时序分类(CTC)目标。这些结果为开发和评估汉语多方言语音处理系统提供了实用指导。
英文摘要
This paper summarizes the ChinaVoices Challenge 2026, which aims to establish unified task definitions and evaluation conditions for Chinese dialect speech processing and to advance multi-dialect identification and automatic speech recognition. The challenge covers 16 dialect categories and defines two tasks: Chinese Multi-Dialect Identification and Chinese Multi-Dialect Automatic Speech Recognition (ASR). It uses approximately 320 hours of speech across the Reference Set, Open Evaluation Set, and Hidden Evaluation Set. The two tasks use the same evaluation audio, and each includes restricted-data and open-data tracks. We describe the task settings, data, evaluation metrics, and Qwen3-ASR-1.7B baseline, and analyze the leaderboard results and submitted systems. In total, 28 teams submit results, 17 provide system reports, and systems from 15 teams pass the compliance review and are included in the analysis. Most eligible systems outperform the baseline, and the official top-three order remains unchanged on the Hidden Evaluation Set for both tasks. Dialect-level results show that categories with higher identification accuracy generally have lower ASR error rates, although the tasks assess related but distinct capabilities. Leading identification systems commonly exploit dialect-discriminative acoustic representations, whereas leading ASR systems emphasize data normalization, augmentation, and auxiliary CTC objectives. These results provide practical guidance for developing and evaluating Chinese multi-dialect speech processing systems.
Comments15 pages,1 figure