发表机构
School of Artificial Intelligence, University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; Wuhan AI Research; GWM AI Lab(中国科学院大学人工智能学院; 中国科学院自动化研究所; 武汉人工智能研究院; 长城汽车人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对低资源汉语方言语音数据稀缺及语义不一致问题,提出DialectS2S模型,通过合成流水线与自对齐后训练策略提升方言语音生成质量,开源框架实现性能突破。
AI 中文摘要
当前端到端语音对话模型主要针对主流语言优化,由于方言语音数据稀缺,在低资源方言场景中仍存在局限。此外,在方言适配过程中,语音对话模型的语义表示空间会持续演变,而传统语音监督保持不变,导致隐藏表示与语音目标之间出现语义不一致,降低了语音的稳定性和自然度。为解决这些问题,我们提出DialectS2S,这是一种面向汉语方言的端到端语音对话模型。我们首先开发了可扩展的方言语音对话合成流水线以高效构建数据,进一步引入带有自对齐语音监督的两阶段后训练策略,该策略将语音监督的语义内容与模型演变后的语义表示对齐,以提升方言语音生成质量。实验结果表明,DialectS2S在多种汉语方言的语音对话任务中均优于现有基线模型,在方言一致性、响应质量和语音可懂度方面实现了显著提升。我们的工作为低资源方言场景下的端到端语音对话建模提供了高效且可扩展的解决方案,为推动未来研究与实际应用,我们已完全开源DialectS2S框架,包括模型检查点、训练数据集和微调代码。
英文摘要
Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speech data. Moreover, during dialect adaptation, the semantic representation space of speech dialogue models continuously evolves, while conventional speech supervision remains unchanged, leading to semantic inconsistency between hidden representations and speech targets and degrading speech stability and naturalness. To address these issues, we propose DialectS2S, an end-to-end speech dialogue model for Chinese dialects. We first develop a scalable dialect speech dialogue synthesis pipeline for efficient data construction. We further introduce a two-stage post-training strategy with self-aligned speech supervision, which aligns the semantic content of speech supervision with the evolved semantic representations of the model to improve dialect speech generation quality. Experimental results show that DialectS2S consistently outperforms existing baselines across multiple Chinese dialects in speech dialogue, achieving substantial improvements in dialect consistency, response quality, and speech intelligibility. Our work provides an efficient and scalable solution for end-to-end speech dialogue modeling in low-resource dialect scenarios. To facilitate future research and practical applications, we fully open-source the DialectS2S framework, including model checkpoints, training datasets, and fine-tuning code.