arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12484eess.AScs.SD

DCASE 2026 挑战赛任务6:长音频中的音频时刻检索——概述与元分析

Overview and Meta-Analysis of DCASE 2026 Challenge Task 6: Audio Moment Retrieval from Long Audio

Hokuto Munakata, Tatsuya Komatsu, Keisuke Imoto, Taichi Nishimura, Huang Xie, Tuomas Virtanen

首次发表
浏览论文内容

中文总结 AI 辅助

本文概述 DCASE 2026 任务6(长音频音频时刻检索),提出结合 MS-CLAP 与 DETR 的基线,最佳系统 Recall1@0.7 达 48.59%,约为基线 3.5 倍,表明增强特征提取与检测网络及校准集成可显著提升性能。

中文摘要 AI 辅助

本文概述了声学场景和事件检测与分类(DCASE)2026 挑战赛任务6,即长音频中的音频时刻检索(AMR)。给定一段数分钟长的音频录音和一个自由形式的文本查询,AMR 旨在检索录音中与查询匹配的时间时刻,其中每个时刻由一对开始和结束时间戳表示。该任务需要有效的跨模态对齐和长范围时间建模。我们描述了任务定义、评估指标、开发和评估数据集,以及一个基线系统,该系统结合了预训练的 MS-CLAP 特征提取器和基于检测变换器(DETR)的时刻检测网络。在开发数据上,基于人工标注数据集和合成数据集训练的基线实现了 13.56% 的 Recall1@0.7,表明长音频中的 AMR 仍然是一个具有挑战性的问题。该挑战吸引了 21 个团队,共提交了 59 个系统。三个最佳系统实现了 48.59% 的 Recall1@0.7,约为基线得分的 3.5 倍。结果表明,加强音频-文本特征提取器和时刻检测网络带来了显著的性能提升。此外,排名前三的团队通过应用置信度分数校准或跨不同时间分辨率的特征集成来提升性能。

英文摘要

This paper presents an overview of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 6, Audio Moment Retrieval (AMR) from Long Audio. Given a several-minute-long audio recording and a free-form text query, AMR aims to retrieve temporal moments in the recording that match the query, where each moment is represented by a pair of start and end timestamps. This task requires effective cross-modal alignment and long-range temporal modeling. We describe the task definition, the evaluation metrics, the development and evaluation datasets, and a baseline system that combines a pre-trained MS-CLAP feature extractor with a Detection Transformer (DETR)-based moment-detection network. On the development data, the baseline trained on a manually annotated dataset and a synthetic dataset achieved Recall1@0.7 of 13.56%, indicating that AMR in long audio remains a challenging problem. The challenge attracted 21 teams, which submitted 59 systems in total. The three best systems achieved Recall1@0.7 of 48.59%, roughly 3.5 times the baseline score. The results show that strengthening the audio-text feature extractor and the moment-detection network led to substantial performance improvements. Furthermore, the top three teams boosted performance by applying confidence score calibration or ensembling across different temporal resolutions of features.

发表机构

  • LY Corporation(LY公司)
  • Kyoto University(京都大学)
  • Sony Interactive Entertainment(索尼互动娱乐)
  • University of Helsinki(赫尔辛基大学)
  • Tampere University(坦佩雷大学)

机构由 AI 辅助整理,请以论文原文为准。

↑