发表机构
Sarvam AI; AI4Bharat, IIT Madras(萨尔瓦姆人工智能公司; 印度理工学院马德拉斯分校人工智能促进印度发展实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
介绍涵盖印度22种在册语言的Indic DiarBench基准数据集,含约108小时多说话者音频,经人工校正注释。通过评估领先系统建立基线,该数据集作为开放资源推动印度语言多语言语音技术研究。
AI 中文摘要
在本研究中,我们介绍了Indic DiarBench,这是一个涵盖印度所有22种在册语言的语音分离和自动语音识别基准数据集。该语料库包含约108小时来自近场会议、远场录音和野外音频的自然多说话者音频。所有注释都经过人工校正,并带有时间对齐的说话者转录。该数据集捕捉了印度语音中常见的对话细微差别。为建立联合自动语音识别和语音分离能力的基线,我们评估了包括商业语音API和多模态大语言模型在内的领先系统。Indic DiarBench作为开放资源发布,以推动印度语言的包容性多语言语音技术研究。
英文摘要
In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. This corpus comprises approximately 108 hours of natural multi-speaker audio from near-field meetings, far-field recordings, and in-the-wild audios. All annotations are human-corrected with time-aligned speaker attributed transcriptions. The dataset captures conversational nuance prevalent in Indian speech, such as English code-mixing, dialectal variation, and frequent speaker overlap. To establish a baseline for joint ASR and diarization capabilities we evaluate leading systems including commercial speech APIs and multimodal large language models. Indic DiarBench is released as an open-access resource to advance inclusive, multilingual speech technology research for Indian languages.
Comments5 pages, 2 figures, Interspeech 2026 conference