发表机构
NAVER LABS; Florida State University(NAVER实验室; 佛罗里达州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究为IWSLT 2026共享任务重新实现NAVER LABS系统,采用规定组件,保留三阶段方法,构建合成示例用于微调,其主要模型在EN-ZH语音翻译和英语SQA任务中取得了一定成绩。
AI 中文摘要
我们为IWSLT 2026共享任务(受限条件、短音频轨道)重新实现了NAVER LABS的IWSLT 2025指令跟随管道,将其适配到规定组件:使用SeamlessM4T-v2-large作为语音编码器,Qwen3-4B-Instruct作为LLM主干。保留了三阶段方法:投影仪对齐、仅文本LoRA预训练和多模态合并。还从提供的语料库中构建了100k个跨十种以语音为中心的任务类型的合成指令跟随示例。主要模型在MCIF基准测试中,EN-ZH语音翻译的COMET为0.781,英语SQA的BERTScore-F1为0.346。
英文摘要
We re-implement the NAVER LABS IWSLT 2025 instruction-following pipeline for the IWSLT 2026 Shared Task (constrained condition, short audio track), adapting it to the mandated components: SeamlessM4T-v2-large as the speech encoder and Qwen3-4B-Instruct as the LLM backbone. The three-stage approach projector alignment, text-only LoRA pre-training, and multimodal merging is preserved from the original design. We additionally construct 100k synthetic instruction-following examples across ten speech-centric task types (10k per task) from the provided corpora, suitable for further Stage 3 fine-tuning. Our primary model achieves COMET 0.781 on EN-ZH speech translation and BERTScore-F1 0.346 on English SQA on the MCIF benchmark.