arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23234cs.CV

多模态大语言模型(MLLM)辅助音频引导的视频目标分割:第8届LSVOS挑战赛MeViS音频赛道第3名报告

MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge

Liangtao Shi, Jinxia Xie, Xiantao Hu, Ting Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出一种结合MLLM与SAM的无训练音频引导视频分割框架,分解任务至各阶段并选用适配基础模型,在第8届LSVOS挑战赛MeViS音频赛道获第3名,验证了基础模型用于该任务的有效性。

中文摘要 AI 辅助

本技术报告提出一种音频引导视频目标分割的无训练框架,将多模态大语言模型(MLLM)与基于SAM的分割模型相结合。我们将该任务分解为多个阶段,并为每个阶段确定合适的基础模型。在不引入额外模型训练或任务特定微调的情况下,我们的方法利用MLLM强大的多模态推理能力来建模文本-视觉对应关系,并采用基于SAM的模型生成准确的目标掩码。该框架证明了利用基础模型进行音频引导视频分割的有效性,并在第8届LSVOS挑战赛的MeViS音频赛道中取得了具有竞争力的性能。

英文摘要

In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. We decompose the task into several stages and identify suitable foundation models for each stage. Without introducing additional model training or task-specific fine-tuning, our approach leverages the strong multimodal reasoning capabilities of MLLMs to model text-visual correspondence and employs SAM-based models for accurate object mask generation. The proposed framework demonstrates the effectiveness of leveraging foundation models for audio-guided video segmentation and achieves competitive performance in the MeViS-Audio Track of the 8th LSVOS Challenge.

发表机构

  • Hefei University of Technology(合肥工业大学)
  • Nanjing University of Science and Technology(南京理工大学)
  • Hunan Police College(湖南警察学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑