arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22214cs.CLcs.SD

百融系统在MLC-SLM 2026中的表现:面向多语言对话语音理解的动态问题感知证据路由

The Bairong System for MLC-SLM 2026: Dynamic Question-Aware Evidence Routing for Multilingual Conversational Speech Understanding

  • Bairong, Inc.(百融云创)

机构由 AI 辅助整理,请以论文原文为准。

Shangkun Huang, Junchao Hu, Huan Shen, Guoji Wang, Yingao Wang, Shaosai Li, Wei Zou, Yunzhang Chen

AI总结:

本文提出百融系统,通过动态问题感知证据路由,在对话文本与音频线索间灵活选择证据,有效提升多语言对话口语问答性能,在MLC-SLM 2026任务中取得显著成果。

AI中文摘要:

长篇幅多语言对话式口语问答要求系统在长距离的对话文本语义与稀疏的声学及说话人敏感线索之间取得平衡。我们提出了用于MLC-SLM 2026挑战赛的百融系统,其中,一个说话人日志-自动语音识别(diarization-ASR)前端生成带说话人属性的文本,一个动态证据路由器根据问题构建特定输入以进行答案预测。该路由器并非采用固定的仅文本或仅音频策略,而是从问题和答案选项中推断所需的证据类型和上下文范围,并在完整对话文本上下文、局部音频-文本融合、说话人关联证据以及紧凑的全局声学样本之间进行选择。这种以文本为主干的设计保持了话语上下文的可用性,同时仅在音频能提供互补证据时激活音频。我们的任务1系统在开发集和评估集上分别实现了25.70%和18.44%的tcpMER。对于任务2,最终系统在开发集上达到了94.84%的准确率,比完整文本基线高出1.68个百分点,比最佳以音频为中心的诊断系统高出2.77个百分点。这些结果支持动态问题感知路由作为对话式口语问答的一种有效证据分配策略。

英文摘要:

Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker-sensitive cues. We present the Bairong system for the MLC-SLM 2026 Challenge, where a diarization-ASR front-end produces speaker-attributed transcripts and a dynamic evidence router constructs question-specific inputs for answer prediction. Instead of applying a fixed transcript-only or audio-only policy, the router infers the required evidence type and context scope from the question and answer options, and selects among full transcript context, local audio-text fusion, speaker-linked evidence, and compact global acoustic samples. This transcript-backbone design keeps discourse context available while activating audio only when it provides complementary evidence. Our Task 1 system achieves 25.70% and 18.44% tcpMER on the development and evaluation sets. For Task 2, the final system obtains 94.84% devel?opment accuracy, outperforming the full-transcript baseline by 1.68 points and the best audio-centric diagnostic system by 2.77 points. These results support dynamic question-aware routing as an effective evidence allocation strategy for conversational spoken QA.

补充信息

↑