arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Auto-AEG:用于开放词汇音频事件定位的可扩展数据构建

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang, Dongjie Fu, Boyun Zhang, Yongbo He, Tao Jin

arXiv 2607.04383首次发表:更新:

AI 中文总结

研究开放词汇音频事件定位任务,因数据稀缺受限。提出Auto-AEG可扩展管道,通过自动数据构建和模型微调构建监督,结合合成音频与伪标签训练,提升模型性能。

AI 中文摘要

大型音频-语言模型在声音推理上流畅,但事件定位不准确,经典声音事件检测标签集封闭。开放词汇音频事件定位任务处于两者交叉,却受数据稀缺瓶颈。我们介绍Auto-AEG,一种通过自动数据构建和模型微调来构建监督的可扩展管道,在基准测试中取得良好性能提升。

英文摘要

Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set. At the intersection of these paradigms lies the task of Open-Vocabulary Audio Event Grounding: predicting all time intervals of a target sound event described by an arbitrary natural language query. Progress is bottlenecked by data scarcity: no large-scale resource provides open-vocabulary onset/offset supervision, and manual temporal annotation is prohibitively expensive. To address this, we introduce Auto-AEG, a scalable pipeline that constructs such supervision by automatic data construction and model fine-tuning. It pairs programmatically synthesized clips, which carry placement-exact ground-truth intervals for supervised cold-start, with multi-model pseudo-labels on real-world audio that supply the reward signal for reinforcement learning. Training with this pipeline yields large temporal-localization gains (+73.9% and +23.1% mIoU over zero-shot) on AEGBench, an independent difficulty-stratified benchmark we release, and these gains generalize to held-out SED and other audio grounding benchmarks. Our results show that automatically constructed data, coupled with interval-aware reward design, provides an effective data-side route to expanding the temporal localization capability of LALMs. AEGBench: https://huggingface.co/datasets/zihan-audio/AEGBench

CommentsWork in progress

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑