arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24064cs.CVcs.AI

MarineEVT:通过视觉工具推理推进以事件为中心的海洋视频理解

MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning

Tuan-An To, Yuk-Kwan Wong, Tuan-Anh Vu, Ziqiang Zheng, Sai-Kit Yeung

首次发表
浏览论文内容

中文总结 AI 辅助

针对海洋视频理解面临的挑战,构建MarineEVT数据集,将其理解分解为EVT-R1过程,利用视觉工具驱动模型。实验显示EVT-R1优于多个SOTA VLMs,为海洋相关发展奠定基础并推动VLMs进步。

中文摘要 AI 辅助

近期视觉语言模型(VLMs)在视觉理解方面取得显著成功,但在视频领域因对时间理解的需求及大规模标注视频数据稀缺而性能下降。本文聚焦海洋视频理解,其面临专业知识需求大、关键信息定位与解读困难等挑战。为此精心构建首个以事件为中心的海洋视频理解数据集MarineEVT,包含20K多任务、视频级视觉问答对。基于此将海洋视频理解分解为以事件为中心的视觉工具集成推理过程EVT-R1,利用强大视觉工具驱动模型定位和解读关键信息。实验表明EVT-R1在不同设置下优于11个SOTA VLMs,为生态发现和海洋教育奠定基础,推动VLMs发展以支持可持续海洋视频理解与分析。

英文摘要

Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-quality image-text pairs. However, the performance of VLMs often degrades in the video domain due to the essential need for temporal understanding and the scarcity of large-scale annotated video data. In this work, we focus on marine video understanding, which brings further challenges: first, it requires substantial domain expertise; and video VLMs usually struggle with localizing and interpreting critical information from marine videos, as the informative events are typically sparse, unpredictable, and unevenly distributed. To address these challenges, we carefully curate the first event-centric marine video understanding dataset called MarineEVT, which features 20K multi-task, video-level visual question-answering pairs spanning multiple dimensions of marine understanding and analysis. Meanwhile, based on MarineEVT, we decompose marine video understanding as an Event-centric Visual Tool-integrated Reasoning process EVT-R1 for short, where we leverage powerful visual tools to drive the model to localize and interpret critical information aligned with visual questions and human intent. To demonstrate its effectiveness, we compare EVT-R1 against 11 SOTA VLMs in different settings. EVT-R1 outperforms the top open-source and top commercial models by 5.22 and 11.09, respectively. MarineEVT and EVT-R1 lay the foundation for ecological discovery and marine education, fostering the development of VLMs capable of interpreting marine dynamics, reasoning about ecological interactions, and supporting sustainable ocean video understanding and analysis.

发表机构

  • The Hong Kong University of Science and Technology(香港科技大学)
  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
  • University of Electronic Science and Technology of China(电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑