发表机构
Telefónica Innovación Digital, Scientific Group; Fondazione Bruno Kessler; Consiglio Nazionale delle Ricerche(西班牙电信数字创新,科学组; 布鲁诺·凯斯勒基金会; 意大利国家研究委员会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出三种方法解决多语言MCQA任务:LoRA微调、多模态上下文学习及免训练检索系统,分别取得0.72、0.81和0.68的宏平均准确率,均大幅超越官方基线。
AI 中文摘要
本文详细介绍了 Eloquence 团队在 Interspeech 2026 第二届 MLC-SLM 挑战赛任务 2 中的方法,该任务涉及跨 21 种语言的多语言多项选择题问答(MCQA)。我们探索了三种方法。首先,我们通过 LoRA 微调 Voxtral-Mini-3B 模型,并结合跨语言数据增强、ASR 转录增强和时间戳感知的音频裁剪,在评估第二阶段取得了 0.72 的宏平均准确率。其次,我们将多模态上下文学习(ICL)应用于冻结的 Voxtral-24B 模型,以纠正强烈的标签偏差,达到了 0.81 的最佳结果。第三,一种基于三层语音锚定记忆(结合声学身份、语义内容和知识图谱)的免训练检索系统取得了 0.68 的成绩。所有三种系统均大幅优于官方基线。
英文摘要
This paper details the Eloquence team's approach to Task 2 of the 2nd MLC-SLM challenge at Interspeech 2026, which involves multilingual Multiple-Choice Question Answering (MCQA) across 21 languages. Three approaches are explored. First, we fine-tune Voxtral-Mini-3B via LoRA with cross-lingual data augmentation, ASR transcript augmentation and timestamp-aware audio cropping, achieving 0.72 macro-accuracy on evaluation Phase 2. Second, we apply multimodal in-context learning (ICL) to the frozen Voxtral-24B model to correct a strong label bias, reaching 0.81, our best result. Third, a training-free retrieval system based on a three-layer voice-anchored memory combining acoustic identity, semantic content, and a knowledge graph achieves 0.68. All three systems substantially outperform the official baseline.