发表机构
Indian Institute of Science, Bangalore(印度科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建了基于Project VAANI的VAANI噪声事件时间戳数据集,含印度165个地区105种语言的自发语音及七类带时间戳的重叠噪声标注,可助力抗噪声ASR、SED等任务。
AI 中文摘要
大多数公开的声音事件语料库要么针对通用音频标记优化,要么针对纯净语音分离优化,而能直接为真实场景下的自发语音提供带时间戳的详细噪声标注的语料库相对较少。我们提出了VAANI噪声事件时间戳数据集,这是一个基于Project VAANI的衍生标注层,Project VAANI包含在印度165个地区用105种语言收集的自发语音实地录音。与合成混合语料库不同,VAANI同时捕获现场的语音和环境噪声,并为每条录音的重叠背景噪声事件标注了精细的起止时间戳,这些噪声事件被归类为紧凑的七类语义分类:动物、交通、婴儿/儿童、音乐、信号/警报、电器和非语音人声。这种自发的多语言印度语音、真实的区域声景以及可与语音目标重叠的跨度级噪声标注的组合,解决了现有数据集仅部分处理的任务:抗噪声自动语音识别(ASR)、声音事件检测(SED)和语音增强。我们将VAANI与九个广泛使用的语料库和基准进行了对比,包括WHAM!、AVA-Speech、MUSAN、FSD50K、CHiME-6、AudioSet、DESED、针对印度的iNoise噪声数据库以及Kathbath-Noisy噪声ASR基准,并描述了用于生成带时间戳标注的标注协议和质量控制流程。
英文摘要
Most public sound-event corpora are optimized either for general audio tagging or for clean speech separation, and comparatively few provide strong timestamped noise annotations layered directly on top of spontaneous, real-world speech. We present the VAANI Noise Event Timestamp Dataset, a derived annotation layer built on Project VAANI field recordings of spontaneous speech collected across 165 Indian districts in 105 languages. Unlike synthetically mixed corpora, VAANI captures speech and ambient noise in situ and simultaneously, and annotates each recording with fine-grained start/end timestamps for overlapping background noise events organized into a compact seven-class semantic taxonomy: animal, traffic, baby/child, music, signal/alarm, appliance, and non-speech human. This combination of spontaneous multilingual Indic speech, authentic regional soundscapes, and span-level noise tags that may overlap with speech targets tasks that existing datasets address only partially: noise-robust Automatic Speech Recognition (ASR), sound event detection (SED), and speech enhancement. We position VAANI against nine widely used corpora and benchmarks, including WHAM!, AVA-Speech, MUSAN, FSD50K, CHiME-6, AudioSet, DESED, the India-specific iNoise noise database, and the Kathbath-Noisy noisy-ASR benchmarks, and describe the annotation protocol and quality-control procedure used to produce the timestamped tags.