AI 中文总结
该研究针对多语言双说话者对话语音,设计了结合语音分离前端与适配Qwen3-ASR-1.7B识别器的系统。通过多步微调及强化学习适配ASR模型,在挑战赛中取得较好成绩,消融实验显示各部分对提升性能和鲁棒性的作用。
AI 中文摘要
本文介绍了我们为2026年多语言双说话者对话语音的MLC-SLM挑战赛任务1自行设计的系统。该系统将模块化的语音分离前端与经过挑战赛适配的Qwen3-ASR-1.7B识别器相结合。语音分离前端执行语音活动检测、子段生成、CAMPPlus说话者嵌入提取、双说话者频谱聚类以及基于RTTM的音频分割。得到的按说话者归属的片段按语言或区域分组,并由适配的ASR模型解码。对于ASR适配,先在官方训练数据上进行有监督的全量微调,然后使用基于三管道TTS的合成语音增强框架生成的合成语音进行LoRA微调,最后使用基于WER/CER的奖励以及对幻觉、重复和长度偏差的惩罚的GRPO强化学习来优化模型。在官方开发集上,完整系统的平均tcpMER为23.70,相对于发布的Qwen-ASR-1.7B性能,错误率绝对降低了6.83个百分点。在最终评估集上,系统的平均tcpMER为17.97。消融结果表明,有监督微调带来的增益最大,而合成语音LoRA适配和强化学习进一步提高了鲁棒性。
英文摘要
This paper describes our self-designed system for Task 1 of the MLC-SLM 2026 Challenge for multilingual two-speaker conversational speech. The system combines a modular speaker diarization front end with a challenge-adapted Qwen3-ASR-1.7B recognizer. The diarization front end performs voice activity detection, subsegment generation, CAMPPlus speaker embedding extraction, two-speaker spectral clustering, and RTTM-based audio segmentation. The resulting speaker-attributed segments are grouped by language or region and decoded by the adapted ASR model. For ASR adaptation, we first perform supervised full fine-tuning on the official training data, then apply LoRA fine-tuning with synthetic speech generated by a three-pipeline TTS-based synthetic speech augmentation framework, and finally refine the model using GRPO reinforcement learning with rewards based on WER/CER and penalties for hallucination, repetition, and length deviation. On the official development set, the full system achieves an average tcpMER of 23.70, reducing the error rate by 6.83 absolute points relative to the released Qwen-ASR-1.7B performance. On the final evaluation set, the system achieves an average tcpMER of 17.97. Ablation results show that supervised fine-tuning provides the largest gain, while synthetic-speech LoRA adaptation and reinforcement learning further improve robustness.
CommentsAccepted by Interspeech 2026 MLC-SLM Workshop