arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DiaScriber:面向多说话人场景的联合说话人 diarization(语音分割)与转录的语音大语言模型

DiaScriber: A Speech LLM for Joint Diarization and Transcription in Multi-Speaker Scenarios

Bingshen Mu, Xian Shi, Xiong Wang, Zhifang Guo, Ting He, Xize Cheng, Yu Xi, Jin Xu, Lei Xie

arXiv 2608.22796首次发表:更新:

发表机构

Northwestern Polytechnical University(西北工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于Qwen3.5-Omni的端到端多说话人 diarization与转录模型DiaScriber,通过构建多样数据流程与三阶段训练策略,在多说话人场景测试集上性能优于对比方法且泛化能力出色。

AI 中文摘要

多说话人自动语音识别(MSASR)旨在联合预测内容转录文本、说话人身份及时间戳,从而解决“谁在何时说了什么”的核心问题,在真实多说话人场景中具有重要实用价值。然而,在快速话轮转换、重叠语音以及复杂多样的多说话人场景下,MSASR仍面临诸多挑战。本研究提出DiaScriber,一种基于语音大语言模型构建的端到端多说话人 diarization(语音分割)与转录模型。我们首先构建多样数据流程以覆盖各类多说话人场景及其复杂性,包括验证与优化、话轮转换与重叠语音模拟,以及多模态标注。此外,DiaScriber基于预训练版本的Qwen3.5-Omni,通过持续预训练、监督微调、强化学习的三阶段训练策略开发。实验表明,DiaScriber在广泛的多说话人场景测试集上取得了优于对比方法的性能,并在未见过的多说话人场景中展现出出色的泛化能力。

英文摘要

Multi-speaker automatic speech recognition (MSASR) aims to jointly predict content transcriptions, speaker identities, and timestamps, thereby addressing the key question of "who spoke what and when" and holds substantial practical value in real-world multi-speaker scenarios. However, MSASR still encounters considerable challenges in the presence of fast turn transitions, overlapping speech, and complex, diverse multi-speaker scenarios. In this work, we propose DiaScriber, an end-to-end multi-speaker diarization and transcription model built on a speech large language model. We first construct diverse data pipelines to cover a wide variety of multi-speaker scenarios and their complexities, including validation and refinement, turn-transition and overlapping-speech simulation, and multimodal annotation. Furthermore, DiaScriber is developed based on the pretrained version of Qwen3.5-Omni through a three-stage training strategy involving continual pretraining, supervised fine-tuning, and reinforcement learning. Experiments show that DiaScriber achieves superior performance over comparison methods across extensive multi-speaker scenario test sets and demonstrates outstanding generalization ability in unseen multi-speaker scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑