多模态大语言模型(MLLM)辅助音频引导的视频目标分割:第8届LSVOS挑战赛MeViS音频赛道第3名报告
MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge
浏览论文内容
中文总结 AI 辅助
该研究提出一种结合MLLM与SAM的无训练音频引导视频分割框架,分解任务至各阶段并选用适配基础模型,在第8届LSVOS挑战赛MeViS音频赛道获第3名,验证了基础模型用于该任务的有效性。
中文摘要 AI 辅助
本技术报告提出一种音频引导视频目标分割的无训练框架,将多模态大语言模型(MLLM)与基于SAM的分割模型相结合。我们将该任务分解为多个阶段,并为每个阶段确定合适的基础模型。在不引入额外模型训练或任务特定微调的情况下,我们的方法利用MLLM强大的多模态推理能力来建模文本-视觉对应关系,并采用基于SAM的模型生成准确的目标掩码。该框架证明了利用基础模型进行音频引导视频分割的有效性,并在第8届LSVOS挑战赛的MeViS音频赛道中取得了具有竞争力的性能。
英文摘要
In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. We decompose the task into several stages and identify suitable foundation models for each stage. Without introducing additional model training or task-specific fine-tuning, our approach leverages the strong multimodal reasoning capabilities of MLLMs to model text-visual correspondence and employs SAM-based models for accurate object mask generation. The proposed framework demonstrates the effectiveness of leveraging foundation models for audio-guided video segmentation and achieves competitive performance in the MeViS-Audio Track of the 8th LSVOS Challenge.
发表机构
- Hefei University of Technology(合肥工业大学)
- Nanjing University of Science and Technology(南京理工大学)
- Hunan Police College(湖南警察学院)
机构由 AI 辅助整理,请以论文原文为准。