arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Indic DiarBench:印度语言的多语言联合语音分离与自动语音识别基准

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

Deovrat Mehendale, Aditya Mehndiratta, Dhruv Rathi, Kaushal Bhogale, Mitesh M. Khapra

arXiv 2607.23808首次发表:更新:

发表机构

Sarvam AI; AI4Bharat, IIT Madras(萨尔瓦姆人工智能公司; 印度理工学院马德拉斯分校人工智能促进印度发展实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

介绍涵盖印度22种在册语言的Indic DiarBench基准数据集,含约108小时多说话者音频,经人工校正注释。通过评估领先系统建立基线,该数据集作为开放资源推动印度语言多语言语音技术研究。

AI 中文摘要

在本研究中,我们介绍了Indic DiarBench,这是一个涵盖印度所有22种在册语言的语音分离和自动语音识别基准数据集。该语料库包含约108小时来自近场会议、远场录音和野外音频的自然多说话者音频。所有注释都经过人工校正,并带有时间对齐的说话者转录。该数据集捕捉了印度语音中常见的对话细微差别。为建立联合自动语音识别和语音分离能力的基线,我们评估了包括商业语音API和多模态大语言模型在内的领先系统。Indic DiarBench作为开放资源发布,以推动印度语言的包容性多语言语音技术研究。

英文摘要

In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field recordings, and in-the-wild audios. All annotations are human-corrected with time-aligned speaker attributed transcriptions. The dataset captures conversational nuance prevalent in Indian speech, such as English code-mixing, dialectal variation, and frequent speaker overlap. To establish a baseline for joint ASR and diarization capabilities we evaluate leading systems including commercial speech APIs and multimodal large language models. Indic DiarBench is released as an open-access resource to advance inclusive, multilingual speech technology research for Indian languages.

Comments5 pages, 2 figures, Interspeech 2026 conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑