基于SAM3的关键帧推理:第8届LSVOS挑战赛MeViS-Text赛道第三名解决方案
Key-Frame Reasoning with SAM3: Third Place Solution for the MeViS-Text Track of the 8th LSVOS Challenge
- Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
- Nanyang Technological University(南洋理工大学)
- Shenzhen Loop Area Institute(深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出两阶段无训练方案,依托Gemini-3.1 Pro分解视频事件选关键帧,用SAM3-agent生成掩码并双向传播,获第8届LSVOS挑战赛MeViS-Text赛道第三名。
AI中文摘要:
本报告提出了一种用于第8届LSVOS挑战赛MeViS-Text赛道的两阶段无训练解决方案。该任务要求模型在整个视频中定位并分割由自然语言表达式指定的对象,此类表达式通常依赖时间线索,包括动作、交互、方向和相对位置。第一阶段通过API使用Gemini-3.1 Pro将视频级事件分解为实例级目标,为每个目标选择一个关键帧,并生成与该帧对齐的判别性描述。第二阶段,SAM3-agent在所选帧上生成像素级种子掩码,SAM3视频跟踪器将该掩码双向传播至整个视频。有效实例在逐帧掩码合并前被独立定位和传播。所有本地SAM3处理均在单块NVIDIA GeForce RTX 4090上运行,无需任务特定训练或模型集成。本方法在挑战赛测试集上排名第三,获得J&F、J、F、N-acc.、T-acc.和最终得分分别为0.761、0.7367、0.7852、0.8333、0.9755和0.856593。
英文摘要:
This report presents a two-stage, training-free solution for the MeViS-Text track of the 8th LSVOS Challenge. The task requires a model to localize and segment the object specified by a natural-language expression throughout a video. Such expressions often depend on temporal cues, including actions, interactions, directions, and relative positions. Our first stage uses Gemini-3.1 Pro via API to decompose a video-level event into instance-level targets, select a key frame for each target, and generate a discriminative description aligned with that frame. In the second stage, SAM3-agent produces a pixel-level seed mask on the selected frame, and the SAM3 video tracker propagates the mask bidirectionally through the video. Valid instances are grounded and propagated independently before their frame-wise masks are merged. All local SAM3 processing runs on a single NVIDIA GeForce RTX 4090 without task-specific training or model ensembling. Our method ranked third on the challenge test set, obtaining J&F, J, F, N-acc., T-acc., and Final scores of 0.761, 0.7367, 0.7852, 0.8333, 0.9755, and 0.856593, respectively.