arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AVSD-Scenes:城市场景音视频描述数据集

AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes

Dhanunjaya Varma Devalraju, Arshdeep Singh, Mark D. Plumbley

arXiv 2610.01861首次发表:更新:

发表机构

IIT Madras(印度理工学院马德拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出AVSD-Scenes,一个包含12,291条配对音视频描述的城市场景数据集,利用多模态大模型生成描述,并通过实验证明多模态描述提升语义对齐和检索性能,场景分类准确率达95.4%。

AI 中文摘要

自然语言描述可以为音视频城市场景提供丰富的语义表示,然而同时描述听觉和视觉信息的数据集仍然有限。在本文中,我们介绍了AVSD-Scenes,一个面向城市环境的配对音视频场景描述数据集。该数据集包含从TAU Urban Audio-Visual Scenes数据集中生成的12,291条音视频场景描述。为了构建该数据集,我们首先分别使用Qwen2-Audio-7B和Qwen2.5-VL-7B生成基于音频和基于视觉的描述。然后,这些特定模态的描述通过大型语言模型(即Qwen3-14B、Mistral-Small-3.2-24B-Instruct-2506和Gemma-3-27B-it)进行组合,以生成捕捉两种模态互补信息的多模态描述。我们使用语义对齐、跨模态检索、场景分类、LLM-as-a-judge评估和人类主观评估对AVSD-Scenes进行基准测试。结果表明,与特定模态的描述相比,多模态描述提高了语义对齐和跨模态检索性能,同时保留了强场景判别信息。生成的描述在城市场景分类中达到了高达94.5%的准确率,而结合音频、视觉和描述嵌入进一步将准确率提高到95.4%。此外,即使从提示指令中移除场景标签,描述仍然具有高度的场景判别性,这表明它们捕捉了源自音视频内容的语义信息,而不仅仅是反映标签信息。

英文摘要

Natural language descriptions can provide rich semantic representations of audio-visual urban scenes, yet datasets that jointly describe both auditory and visual information remain limited. In this paper, we introduce AVSD-Scenes, a paired audio-visual scene description dataset for urban environments. The dataset contains 12,291 audio-visual scene descriptions generated from the TAU Urban Audio-Visual Scenes dataset. To construct the dataset, we first generate audio- and visual-based descriptions using Qwen2-Audio-7B and Qwen2.5-VL-7B, respectively. These modality-specific descriptions are then combined using large language models, namely Qwen3-14B, Mistral-Small-3.2-24B-Instruct-2506, and Gemma-3-27B-it, to produce multimodal descriptions that capture complementary information from both modalities. We benchmark AVSD-Scenes using semantic alignment, cross-modal retrieval, scene classification, LLM-as-a-judge evaluation, and human subjective assessment. Results show that multimodal descriptions improve semantic alignment and cross-modal retrieval performance compared with modality-specific descriptions while preserving strong scene-discriminative information. The generated descriptions achieve up to 94.5% accuracy in urban scene classification, while combining audio, visual, and description embeddings further improves accuracy to 95.4%. Furthermore, the descriptions remain highly scene-discriminative even when scene labels are removed from the prompting instructions, indicating that they capture semantic information derived from the audio-visual content rather than merely reflecting label information.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑