arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.02269cs.CVcs.AI

AnyGroundBench: 视觉语言模型中视频定位的专业领域基准

AnyGroundBench: A Multi-Domain Adaptation Benchmark for Video Grounding in VLMs

Rintaro Otsubo, Ryo Fujii, Reina Ishikawa, Taiki Kanaya, Kanta Sawafuji, Hiroki Kajita, Shigeki Sakai, Hideo Saito, Ryo Hachiuma

首次发表
浏览论文内容

中文总结 AI 辅助

提出AnyGroundBench基准,将STVG评估从零样本转向领域适应,涵盖五个专业领域,评估15个VLM的零样本和上下文学习能力,发现现有模型在专业领域表现不佳。

中文摘要 AI 辅助

视觉语言模型(VLM)在时空视频定位(STVG)中展现出巨大潜力。然而,当前的评估协议主要局限于通用日常基准上的零样本评估。这与专业领域的实际应用存在严重脱节,因为模型不可避免地会遇到罕见的视觉概念和复杂的时空动态。由于跨无限数据分布的全面预训练不可行,适应新领域的能力至关重要。为弥合这一差距,我们引入了AnyGroundBench,一个领域适应基准,旨在将STVG评估范式从静态零样本测试转变为严格的领域适应。针对五个专业领域(动物、工业、体育、手术和公共安全),AnyGroundBench将新拍摄的视频(如专家标注的小鼠行为)与现有数据集配对,通过密集、高保真的时空标注统一它们。关键的是,该基准提供了专门的训练子集,以系统衡量领域适应性。我们广泛评估了15个最先进的VLM,评估了它们在实际计算约束下的零样本泛化和上下文学习(ICL)能力。最终,我们的发现揭示了当前模型在面对专业领域时,在零样本和基于ICL的适应中均失败,暴露了未来研究必须解决的时空推理中的关键缺陷。

英文摘要

Vision-Language Models (VLMs) have shown strong performance in Spatio-Temporal Video Grounding (STVG), yet they are still evaluated mostly in a zero-shot manner on general-purpose benchmarks of everyday scenes. This creates a critical disconnect from real-world applications in specialized domains, where models inevitably encounter rare visual or textual concepts. Since exhaustive pre-training across infinite data distributions is infeasible, the ability to adapt to novel domains with limited data is essential. To bridge this gap, we introduce AnyGroundBench, a domain-adaptation benchmark designed to shift the STVG evaluation paradigm from static zero-shot testing to rigorous domain adaptation. Targeting five specialized domains (animal, industry, sports, surgery, and public security), AnyGroundBench pairs newly captured, expert-annotated videos with established datasets, unifying them through dense, high-fidelity spatio-temporal annotations. Crucially, the benchmark provides dedicated limited training subsets, enabling systematic evaluation of domain adaptability under limited training data. We benchmark 23 state-of-the-art VLMs in the zero-shot setting and further evaluate five adaptation strategies, spanning training-free and fine-tuning-based approaches, on representative models, assessing their zero-shot generalization and adaptation capacity. Our results show that current VLMs remain far from practical performance in the zero-shot setting, that training-free adaptation produces highly variable effects depending on the model and domain, and that fine-tuning-based adaptation, though more effective, still falls short of real-world requirements, with gains varying markedly across domains. These findings expose fundamental limitations in current VLMs' spatio-temporal reasoning, pointing to concrete directions for future research.

发表机构

  • Keio University(庆应大学)
  • Keio AI Research Center(庆应人工智能研究中心)
  • Keio University School of Medicine(庆应大学医学部)
  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

↑