arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SEAR:面向弱到强多语言语音多项选择题的片段证据感知路由

SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ

Huy Hoang Le, Long-Bao Nguyen, Minh Tri Dao

arXiv 2609.11355首次发表:更新:

发表机构

CAKE by VPBank(VPBank旗下CAKE)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SEAR方法,通过片段证据感知的数据生成与路由,将弱文本可答和强音频依赖的MCQ分别用于监督微调和强化学习,在MLC-SLM挑战赛任务2中达到90.92%准确率。

AI 中文摘要

本文描述了我们在第二届多语言对话语音语言模型(MLC-SLM)挑战赛任务2中的系统。我们采用片段证据感知的数据和后训练流程来适配Qwen3-Omni-30B-A3B-Instruct模型。一个语言模型将带时间戳的自动语音识别(ASR)结果转换为连贯的事件片段,这些片段通过边界余量进行扩展,并从原始录音中裁剪出来。随后,我们使用Qwen3.6-27B合成互补的语义多项选择题(MCQs),并使用Gemini 3.1 Flash-Lite合成声学多项选择题,接着进行结构、接地、答案一致性和目标模型可训练性检查,最终在21种语言和口音变体中生成359,825个经过验证的多项选择题。一个仅文本的探针将数据划分为弱项(文本可回答)用于监督微调,以及强项(依赖音频)用于通过组序列策略优化(GSPO)进行强化学习,并通过去偏优势、序列级重要性校正和动态过滤来稳定训练。我们的系统在最终官方评估集上达到了90.92%的准确率。

英文摘要

This paper describes our system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language model converts timestamped ASR into coherent event spans, which are expanded by a boundary margin and cropped from the original recording. We then synthesize complementary semantic MCQs with Qwen3.6-27B and acoustic MCQs with Gemini~3.1 Flash-Lite, followed by structural, grounding, answer-consistency, and target-model trainability checks, yielding 359,825 verified MCQs across 21 language and accent variants. A text-only probe partitions the data into weak, text-answerable items used for supervised fine-tuning and strong, audio-dependent items used for reinforcement learning with Group Sequence Policy Optimization (GSPO), stabilized by debiased advantages, sequence-level importance correction, and dynamic filtering. Our system obtains 90.92% accuracy on the final official evaluation set.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑