arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多智能体目标存在性验证与学习型掩码几何优化:2026年第8届LSVOS挑战赛MeViS-Text赛道获胜报告

Multi-Agent Target-Existence Verification and Learned Mask Geometry Refinement: Winning Report of the MeViS-Text Track at the 8th LSVOS Challenge 2026

Jungyoon Lee, Gyuil Lim, Doeon Kim, Seong-heum Kim

arXiv 2608.11458首次发表:更新:

发表机构

Soongsil University(崇实大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出SSUPER方案,通过多智能体审计解耦存在性验证、StyleRefiner优化掩码几何,获2026年LSVOS挑战赛MeViS-Text赛道冠军,最终得分0.9081339614。

AI 中文摘要

我们提出了2026年第8届大规模视频对象分割(LSVOS)挑战赛MeViS-Text赛道的第一名解决方案:该任务为基于书面运动表达式的 referring视频对象分割,包含欺骗性无目标表达式,此类表达式与视频中任何对象均不匹配,必须在每一帧中生成空掩码。我们的流水线SSUPER将每个表达式解析为视觉概念,使用SAM 3.1生成全视频候选掩码片段,并选择目标ID。在每个推理阶段,三个异构多模态大语言模型独立执行相同的阶段特定提示,之后通过一次合成过程输出经模式验证的裁决。尽管该系统在验证阶段拒绝了所有无目标表达式,但排行榜显示仍有大量测试无目标案例未被检测到。原因在于,困难负样本命名了合理对象,仅在完整时间谓词下才不匹配,因此当选择与存在性判定结合进行时,类别合理的掩码片段会锚定裁决。因此,我们将存在性验证解耦为对完整谓词(类别、数量、动作、轨迹、事件顺序和语义角色)的独立多智能体审计,该审计可区分不存在与临时不可见,抵消相机运动导致的表观运动,并要求矛盾证据而非仅不确定性来做出无目标裁决。无需任何新的分割调用,该审计即可纠正大部分剩余的无目标错误。仅使用训练数据的StyleRefiner随后将掩码几何与MeViSv2的标注风格对齐,同时通过构造保留所有存在判定,表明一旦语义固定,剩余部分错误中部分属于风格而非语义错误。完整系统在官方挑战排行榜上达到了0.9081339614的最终分数。

英文摘要

We present the first-place solution to the MeViS-Text track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge 2026: referring video object segmentation guided by written motion expressions, including deceptive no-target expressions that match no object in the video and must yield empty masks in every frame. Our pipeline, SSUPER, resolves each expression into a visual concept, generates full-video candidate masklets with SAM~3.1, and selects target IDs. At every reasoning stage, three heterogeneous multimodal large language models independently execute the same stage-specific prompt before a single synthesis pass commits one schema-validated verdict. Although this system rejects every no-target expression in validation, the leaderboard reveals that a substantial share of test no-target cases still slips through. The reason is that hard negatives name a plausible object and fail only under the complete temporal predicate, so when selection and existence are decided together, a category-plausible masklet anchors the verdict. Hence, we decouple existence verification into an independent multi-agent audit of the full predicate (category, count, action, trajectory, event order, and semantic role) that distinguishes absence from temporary invisibility, discounts apparent motion caused by camera movement, and requires contradicting evidence rather than mere uncertainty for a no-target verdict. Without any new segmentation call, this audit recovers most of the residual no-target errors. A training-data-only StyleRefiner then aligns mask geometry with the annotation style of MeViSv2 while preserving every presence decision by construction, showing that once the semantics are fixed, part of the remaining error is stylistic rather than semantic. The complete system reaches a Final score of 0.9081339614 on the official challenge leaderboard.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑