发表机构
Institute of Computational Perception, Johannes Kepler University Linz, Austria; LIT Artificial Intelligence Lab, Linz, Austria(计算感知研究所,约翰尼斯·开普勒大学林茨,奥地利; LIT人工智能实验室,林茨,奥地利)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出RealDESED真实世界家庭声音事件检测基准,由自然环境音频记录构成,具多注释者标注和丰富元数据。建立基于Transformer的基线并研究多种策略方法,该基准有助于弥合当前研究基准与实际部署间差距。
AI 中文摘要
本文介绍了RealDESED,一个真实世界的家庭声音事件检测(SED)基准,包含由652名参与者在其家中收集的5710段音频记录。每段记录时长15至35秒,包含15种常见家庭声音类别的精确时间注释。与现有SED数据集不同,RealDESED完全由自然家庭环境中捕获的记录组成。其具有多注释者标注方案,数据集还提供丰富元数据。建立了基于Transformer的基线并研究相关策略和方法,基线在测试集上的宏平均PSDS1分数为0.731。RealDESED为开发和评估鲁棒SED系统提供了有价值的基准。
英文摘要
This paper presents RealDESED, a real-world domestic sound event detection (SED) benchmark comprising 5,710 audio recordings collected by 652 participants in their homes. Each recording is between 15 and 35 seconds long and contains temporally precise annotations for 15 common domestic sound classes. In contrast to existing SED datasets, which typically rely on simulated soundscapes or broad web-crawled audio, RealDESED consists exclusively of recordings captured in natural domestic environments, reflecting realistic variability in recording devices, device placement, acoustic conditions, background sounds, and naturally occurring event co-occurrences. A distinguishing characteristic of the dataset is its multi-annotator labeling scheme, where each recording is independently annotated by multiple annotators, while the validation and test sets undergo an additional review process to ensure high annotation quality and reliable benchmarking. Furthermore, the dataset provides rich metadata, including recording device, device placement, environment labels, and textual scene descriptions. We establish a strong transformer-based baseline and investigate annotation aggregation strategies, post-processing methods, long-form inference, and the impact of recording metadata on model performance. Our baseline achieves a macro-averaged PSDS1 score of 0.731 on the test set. We believe RealDESED provides a valuable benchmark for developing and evaluating robust SED systems under realistic domestic conditions, helping to bridge the gap between current research benchmarks and real-world deployment.
CommentsSubmitted to the DCASE 2026 Workshop (Detection and Classification of Acoustic Scenes and Events). Resources: Dataset (Zenodo): https://zenodo.org/records/20056072; code and baseline implementation (GitHub): https://github.com/fschmid56/RealDESED