发表机构
Shenzhen Transsion Holdings Co., Ltd(深圳传音控股股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出由说话人日志、多语言ASR和融合模块组成的级联框架,解决多语言对话语音的说话人属性转录问题,在MLC-SLM 2026挑战赛任务1中取得15.41%的tcpMER并排名第二。
AI 中文摘要
本文介绍了传音语音团队对MLC-SLM 2026挑战赛任务1的提交方案,该任务聚焦于多语言对话语音的说话人属性转录。我们提出了一种由三个组件组成的级联框架:说话人日志模块、长形式多语言ASR模块以及说话人-转录融合模块。日志模块基于DiariZen构建,通过局部说话人活动估计和全局说话人聚类生成说话人同质片段。ASR模块基于Qwen3-Omni,生成多语言转录,同时外部基于CTC的对齐模型提供精确的词级和字符级时间戳。最后,融合模块将日志输出与带时间戳的转录相结合,生成带说话人属性的STM输出。在官方评估集上的实验结果表明了所提出框架的有效性。提交的系统实现了15.41%的tcpMER,在所有参赛队伍中排名第二。
英文摘要
This paper presents the Transsion Speech Team submission to Task 1 of the MLC-SLM 2026 Challenge, which focuses on speaker-attributed transcription for multilingual conversational speech. We propose a cascaded framework consisting of three components: a speaker diarization module, a long-form multilingual ASR module, and a speaker-transcription fusion module. The diarization module is built upon DiariZen and produces speaker-homogeneous segments through local speaker activity estimation and global speaker clustering. The ASR module is based on Qwen3-Omni and generates multilingual transcriptions, while an external CTC-based alignment model provides precise word- and character-level timestamps. Finally, the fusion module combines diarization outputs with timestamped transcriptions to generate speaker-attributed STM outputs. Experimental results on the official evaluation set demonstrate the effectiveness of the proposed framework. The submitted system achieves a tcpMER of 15.41% and ranks second among all participating teams.