arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向第二届MLC-SLM挑战赛的前置静音增强与多阶段合成监督

Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge

Kexin Shi, Renhe Sun, Yuge Huang, Ximeng Wang, Jiayi Zhou, Jian Liu, Malu Zhang

arXiv 2608.14150首次发表:更新:

发表机构

Ant Group; UESTC(蚂蚁集团; 电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对第二届MLC-SLM挑战赛的两项任务,分别采用前置静音增强等策略微调VibeVoice-ASR-7B、用合成问答对微调Qwen3-Omni-30B-A3B-Instruct,提升了任务1和任务2的评估性能。

AI 中文摘要

第二届多语种会话语音语言模型(MLC-SLM)挑战赛针对完整、未分段的多语种会话评估两项任务:说话人 diarization(说话人分割)与识别(任务1)、会话语音理解(任务2)。两项任务在评估时均不提供 oracle 话语边界或说话人标签,且任务2无问答训练集。针对任务1,我们通过随机前置静音裁剪、一致时间戳修正及指数移动平均(EMA)训练策略微调VibeVoice-ASR-7B;针对任务2,我们通过多模态候选生成、静音音频过滤及分布匹配增强构建合成问答对,并微调Qwen3-Omni-30B-A3B-Instruct以实现带标签的直接回答。在任务1评估集上,裁剪将tcpMER从18.30%降至17.27%,EMA进一步将其降至16.73%;在任务2评估集上,联合应用分布匹配增强与带标签直接回答将准确率从83.0%提升至86.0%。

英文摘要

The second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge evaluates two tasks over complete, unsegmented multilingual conversations: speaker diarization and recognition (Task 1) and conversational speech understanding (Task 2). Neither task provides oracle utterance boundaries or speaker labels at evaluation, and Task 2 provides no question-answer training set. For Task 1, we fine-tune VibeVoice-ASR-7B with random leading-silence cropping, consistent timestamp correction, and an exponential moving average (EMA) training strategy. For Task 2, we construct synthetic question-answer pairs through multimodal candidate generation, silent-audio filtering, and distribution-matched augmentation, and fine-tune Qwen3-Omni-30B-A3B-Instruct for tagged direct answering. On the Task 1 evaluation set, cropping reduces tcpMER from 18.30% to 17.27%, and EMA further reduces it to 16.73%. On the Task 2 evaluation set, jointly applying distribution-matched augmentation and tagged direct answering raises accuracy from 83.0% to 86.0%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑