发表机构
Nanjing University; Jilin University; Beijing Institute of Technology; National Taiwan University; Shenzhen Pimei Technology Co., Ltd.; Northwestern Polytechnical University(南京大学; 吉林大学; 北京理工大学; 国立台湾大学; 深圳匹美科技有限公司; 西北工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对NVVSpeech挑战赛中非语言发声识别任务,提出基于标签协调与两阶段采样(平方根类别采样+均匀微调)的数据中心ASR流水线,缓解长尾分布,获官方评分63.86并排名第四。
AI 中文摘要
非语言发声(NVVs)携带重要的副语言信息,但常被传统自动语音识别(ASR)系统忽略。ISCSLP NVVSpeech挑战赛要求在有限且高度不平衡的监督下,联合转录词汇内容和16种NVV类别。我们提出了一种基于跨数据集标签协调和两阶段采样调度的数据为中心的NVV感知ASR流水线。我们将异构源标签映射到官方分类体系,并排除无法可靠映射的样本。我们的调度首先使用平方根类别采样来缓解长尾分布,然后应用均匀类别微调。在固定的本地验证集上,平方根类别采样在测试的单阶段设置中表现最佳。最终的两阶段系统获得官方评分63.86,在Track 1中排名第四。
英文摘要
Non-verbal vocalizations (NVVs) carry important paralinguistic information but are often omitted by conventional automatic speech recognition (ASR) systems. The ISCSLP NVVSpeech Challenge requires joint transcription of lexical content and 16 NVV categories under limited and highly imbalanced supervision. We present a data-centric NVV-aware ASR pipeline based on cross-dataset label harmonization and a two-stage sampling schedule. We map heterogeneous source labels to the official taxonomy and exclude samples without a reliable mapping. Our schedule first uses square-root category sampling to moderate the long-tailed distribution and then applies uniform-category fine-tuning. On a fixed local validation split, square-root category sampling performs best among the tested single-stage settings. The final two-stage system obtains an official score of 63.86 and ranks fourth in Track 1.
CommentsAccepted by ISCSLP 2026, NVVSpeech Challenge Track 1