arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HINTT 提交至第二届 MLC-SLM 挑战赛:级联与统一方法在说话人日志和语音识别中的比较

HINTT Submission to the 2nd MLC-SLM Challenge: Comparing Cascaded and Unified Approaches to Diarization and ASR

Takanori Ashihara, Kohei Matsuura, Masato Mimura

arXiv 2610.08063首次发表:更新:

发表机构

NTT, Inc., Japan(日本电信电话公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 HINTT 系统,比较级联(说话人日志+语音识别)与统一语音大语言模型两种方法,实验表明级联系统在 MLC-SLM 任务中更可靠,统一模型是未来方向。

AI 中文摘要

本文介绍了提交至第二届多语言会话语音语言模型挑战赛与研讨会(MLC-SLM)的 HINTT 系统。我们解决多语言说话人属性语音识别问题,即系统必须确定谁在何时说了什么。我们研究了两种建模策略:一种是将说话人日志与基于语音大语言模型的语音识别相结合的级联流水线,另一种是直接生成说话人标签、时间戳和转录文本的统一语音大语言模型。我们的最终提交基于级联流水线,包括微调的 DiariZen 说话人日志模型、微调的 Qwen3-ASR 模型以及基于大语言模型的生成式错误纠正。为了比较,我们还使用相同的官方训练数据微调了 VibeVoice-ASR 作为统一模型。所有任务特定的微调和模型选择仅使用官方 MLC-SLM 数据,不使用外部数据或伪标签。实验结果表明,在 MLC-SLM 任务 1 的条件下,级联系统仍然更可靠,而统一语音大语言模型为未来的说话人属性语音识别提供了一个有前景的方向。

英文摘要

This paper presents the HINTT system submitted to the 2nd Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM). We address multilingual speaker-attributed ASR, where systems must determine who spoke when and what was spoken. We investigate two modeling strategies for this problem: a cascaded pipeline that combines speaker diarization with speech-LLM-based ASR, and a unified speech LLM that directly generates speaker labels, timestamps, and transcriptions. Our final submission is based on the cascaded pipeline, consisting of a fine-tuned DiariZen diarization model, a fine-tuned Qwen3-ASR model, and LLM-based generative error correction. For comparison, we also fine-tune VibeVoice-ASR as a unified model using the same official training data. All task-specific fine-tuning and model selection are performed using only the official MLC-SLM data, without external data or pseudo-labels. Experimental results demonstrate that the cascaded system remains more reliable under the MLC-SLM Task 1 conditions, while unified speech LLMs offer a promising direction for future speaker-attributed ASR.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑