AI 中文总结
该研究针对第二届MLC-SLM挑战赛任务1,提出级联分帧识别系统,结合多种技术,测试无先验信息。分析工程选择影响,如基于嵌入的聚类更优,重叠感知分割虽提召回率但增加tcpMER,在开发集和评估集有相应表现。
AI 中文摘要
我们描述了对第二届MLC-SLM挑战赛任务1的提交:一个级联的先分帧再识别系统,它结合了DiariZen-Large-s80(WavLM-Large)分割、基于CAM++嵌入的双说话人聚类以及一个LoRA适配的全语言自动语音识别语言模型7B v2识别器,测试时无先验分割或说话人标签。在官方开发集(150个对话,21种语言/口音类别)上,该系统的宏tcpMER为29.27%,而官方基线为79.?5%;在评估集上得分为50.23%。我们还分析了两个对tcpMER有重大影响的工程选择。首先,基于嵌入的说话人聚类优于仅从自动语音识别轮次标记分配说话人的端到端方式替代方案。其次,重叠感知分割虽然旨在提高分帧召回率,但由于重叠语音被转录两次,会增加tcpMER。
英文摘要
We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Large-s80 (WavLM-Large) segmentation, CAM++ embedding-based two-speaker clustering, and a LoRA-adapted omniASR LLM 7B v2 recognizer, with no oracle segmentation or speaker labels at test time. On the official Development set (150 conversations, 21 language/accent categories) the system attains a macro tcpMER of 29.27%, versus 79.15% for the official baseline; on the Evaluation set it scores 50.23%. We also analyze two engineering choices that substantially affect tcpMER. First, embedding-based speaker clustering outperforms an end-to-end-style alternative that assigns speakers from ASR <sc> turn markers alone. Second, overlap-aware segmentation, although intended to raise diarization recall, increases tcpMER because overlapped speech is transcribed twice.
CommentsAccepted to INTERSPEECH 2026. 4 pages + references. Technical description of our 2nd MLC-SLM Challenge Task 1 submission