arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当预测空结果胜过 SAM 3:重新审视视频对象分割中的评估

When Predicting Nothing Beats SAM 3: Revisiting Evaluation in Video Object Segmentation

Jihwan Hong, Woohyeon Park, Jaeik Kim, Jaeyoung Do

arXiv 2610.02946首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长视频中目标间歇出现的 VOS 评估缺陷,提出 FaVOS 基准和体积 J&F 指标,以纠正空预测奖励偏差,提升评估与实际性能的一致性。

AI 中文摘要

视频对象分割(VOS)在复杂和长视频中对于现实世界应用日益重要,其中目标对象往往仅在长时间跨度内间歇性出现。然而,现有基准主要关注在视频大部分时间内保持可见的时间显著性对象。为解决这一差距,我们引入了 FaVOS(具有分数时间可见性的视频对象分割基准),这是一个旨在低时间可见性条件下评估 VOS 方法的基准。我们表明,在这种条件下,标准的 J&F 指标可能将 VOS 评估退化为缺席分类,因为空预测在目标缺席帧上获得高奖励。因此,即使是平凡的空掩膜预测器也能胜过如 SAM 3 这样的强模型,揭示了当前指标与实际 VOS 性能之间的根本性不匹配。为缓解此问题,我们提出了体积 J&F,它将掩膜序列作为时空体积进行评估,并减少目标缺席奖励的主导地位,同时保持对分割质量和时间结构的敏感性。项目页面:此 https URL。

英文摘要

Video Object Segmentation (VOS) in complex and long videos is increasingly important for real-world applications, where target objects often appear only intermittently within long temporal horizons. However, existing benchmarks largely focus on temporally salient objects that remain visible for most of the video. To address this gap, we introduce FaVOS (A Benchmark for Video Object Segmentation with Fractional Temporal Visibility), a benchmark designed to evaluate VOS methods under low temporal visibility. We show that, in this regime, the standard J&F metric can collapse VOS evaluation into absence classification, because empty predictions receive high rewards on target-absent frames. Consequently, even a trivial empty-mask predictor can outperform strong models such as SAM 3, revealing a fundamental mismatch between current metrics and practical VOS performance. To mitigate this issue, we propose Volumetric J&F, which evaluates mask sequences as spatio-temporal volumes and reduces the dominance of target-absence rewards while preserving sensitivity to segmentation quality and temporal structure. Project page: https://aidaslab.github.io/FaVOS.

CommentsNeurIPS 2026 E&D

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑