arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视听火烈鸟:用于长而复杂视频的开放视听智能

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze, Nishit Anand, Siddharth Gururani, Hanrong Ye, Pritam Biswas, Yuanhang Su, Ehsan Hosseini-Asl, Sang-gil Lee, Zhifeng Kong, Jaehyeon Kim, Sungwon Kim, S Sakshi, Ramani Duraiswami, Dinesh Manocha, Andrew Tao, Mohammad Shoeybi, Bryan Catanzaro, Ming-Yu Liu, Wei Ping

arXiv 2607.16107首次发表:更新:

发表机构

NVIDIA, USA; University of Maryland, USA(美国NVIDIA公司; 美国马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出视听火烈鸟,用于长复杂视频的联合理解与推理。贡献包括构建大规模视频集、设计三阶段课程及推理框架。实验显示其在多基准测试中表现优异,超越同类开放模型,在长复杂视听理解推理任务中竞争力强,有现实应用价值和泛化能力。

AI 中文摘要

我们提出了视听火烈鸟(AV - Flamingo),这是一种完全开放的最先进的视听大语言模型(AV - LLM),用于对音频、图像和长视频进行联合理解与推理。与之前主要关注短视频片段的AV - LLM不同,AV - Flamingo旨在对长而复杂的现实世界(视听)视频进行理解和推理。为此,我们做出了三项关键贡献:一是视听技能,一个大规模的现实世界视频集合;二是新颖的三阶段课程;三是时间视听交错思维链推理框架。实验表明AV - Flamingo在多个基准测试中表现出色,具有很强的现实应用价值和泛化能力。

英文摘要

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.

CommentsProject Page: https://avflamingo.pages.dev/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑