AI 中文总结
本研究发布了总计2100小时的两个互补中文多方对话语音语料库SmoothConv与DuplexConv,含同步语音轨道与细粒度标注,经SmoothConv基准在多项任务上验证了其用于多方语音交互建模的实用性。
AI 中文摘要
大型音频语言模型(LALMs)的最新进展推动了自然智能语音交互系统的发展,这类系统需要对复杂的对话行为进行建模,包括话轮转换、重叠语音和说话人协调。多方对话为研究这些行为提供了真实场景,但现有的中文对话语料库往往缺乏同步的参与者级语音轨道和全面的标注。本研究引入了SmoothConv和DuplexConv这两个互补的中文多方对话语音语料库,总计2100小时。SmoothConv提供经人工验证的对话,用于可靠的分析与评估;DuplexConv则通过可扩展的流水线提供大规模自动标注的对话,用于模型训练。两个语料库均提供同步的参与者级语音轨道和多维细粒度标注。我们还发布了SmoothConv基准,并在语音分离、多说话人自动语音识别(MSASR)和话轮检测任务上对这些资源进行评估。实验结果表明,所提出的资源对多方语音交互建模具有实用性,相关数据集、基准及资源均已公开。
英文摘要
Recent advances in large audio language models (LALMs) have driven the development of natural and intelligent speech interaction systems. Such systems need to model complex conversational behaviors, including turn-taking, overlapping speech, and speaker coordination. Multi-party conversations provide a realistic setting for studying these behaviors, yet existing Mandarin conversational corpora often lack synchronized participant-level speech tracks and comprehensive annotations. In this work, we introduce SmoothConv and DuplexConv, two complementary Mandarin multi-party conversational speech corpora totaling 2,100 hours. SmoothConv provides human-verified conversations for reliable analysis and evaluation, while DuplexConv offers large-scale automatically annotated conversations through a scalable pipeline for model training. Both corpora provide synchronized participant-level speech tracks and multi-dimensional fine-grained annotations. We further release the SmoothConv Benchmark and evaluate these resources on speech separation, multi-speaker automatic speech recognition (MSASR), and turn detection tasks. Experimental results demonstrate the utility of the proposed resources for multi-party speech interaction modeling. The datasets, benchmark, and related resources are publicly available.