发表机构
Meta(Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对ASR系统忽略非语言发声的问题,提出三种数据驱动策略(两阶段课程学习、跨类别迁移、语音转换增强),利用发声间共享声学结构提升稀有类别识别。
AI 中文摘要
现代自动语音识别(ASR)系统在转录词汇内容方面表现出色,但常常忽略非语言发声(NVs),如笑声、呼吸、咳嗽和哭声,这些发声携带对话和情感信息。在ASR中建模NVs具有挑战性,因为NV标注稀疏且高度长尾,频繁类别如呼吸和笑声主导了罕见事件如哭声和咳嗽。我们研究了三种以数据为中心的策略来改进低资源NV识别:(1)两阶段课程学习,首先将所有NV事件映射到通用标记,然后在目标类别上微调;(2)从高资源事件(如笑声和呼吸)到罕见事件(如哭泣)的跨标记迁移;(3)带类别平衡的语音转换增强。实验表明,可以利用发声事件间的共享声学结构来改善稀有类别检测,同时保持词汇ASR质量。
英文摘要
Modern automatic speech recognition (ASR) systems excel at transcribing lexical content but often omit nonverbal vocalizations (NVs), such as laughter, breaths, coughs, and cries, that carry conversational and affective information. Modeling NVs in ASR is challenging because NV annotations are sparse and highly long-tailed, with frequent categories such as breaths and laughter dominating rarer events such as cries and coughs. We study three data-centric strategies for improving low-resource NV recognition: (1) a two-stage curriculum that first maps all NV events to a generic token and then fine-tunes on target categories; (2) inter-token transfer from high-resource events, such as laughter and breath, to rare events, such as crying; and (3) voice-conversion augmentation with class balancing. Experiments show that shared acoustic structure across vocal events can be exploited to improve rare-category detection while preserving lexical ASR quality.