AnyGroundBench: 视觉语言模型中视频定位的专业领域基准
AnyGroundBench: A Multi-Domain Adaptation Benchmark for Video Grounding in VLMs
浏览论文内容
中文总结 AI 辅助
提出AnyGroundBench基准,将STVG评估从零样本转向领域适应,涵盖五个专业领域,评估15个VLM的零样本和上下文学习能力,发现现有模型在专业领域表现不佳。
中文摘要 AI 辅助
视觉语言模型(VLM)在时空视频定位(STVG)中展现出巨大潜力。然而,当前的评估协议主要局限于通用日常基准上的零样本评估。这与专业领域的实际应用存在严重脱节,因为模型不可避免地会遇到罕见的视觉概念和复杂的时空动态。由于跨无限数据分布的全面预训练不可行,适应新领域的能力至关重要。为弥合这一差距,我们引入了AnyGroundBench,一个领域适应基准,旨在将STVG评估范式从静态零样本测试转变为严格的领域适应。针对五个专业领域(动物、工业、体育、手术和公共安全),AnyGroundBench将新拍摄的视频(如专家标注的小鼠行为)与现有数据集配对,通过密集、高保真的时空标注统一它们。关键的是,该基准提供了专门的训练子集,以系统衡量领域适应性。我们广泛评估了15个最先进的VLM,评估了它们在实际计算约束下的零样本泛化和上下文学习(ICL)能力。最终,我们的发现揭示了当前模型在面对专业领域时,在零样本和基于ICL的适应中均失败,暴露了未来研究必须解决的时空推理中的关键缺陷。
英文摘要
Vision-Language Models (VLMs) have shown strong performance in Spatio-Temporal Video Grounding (STVG), yet they are still evaluated mostly in a zero-shot manner on general-purpose benchmarks of everyday scenes. This creates a critical disconnect from real-world applications in specialized domains, where models inevitably encounter rare visual or textual concepts. Since exhaustive pre-training across infinite data distributions is infeasible, the ability to adapt to novel domains with limited data is essential. To bridge this gap, we introduce AnyGroundBench, a domain-adaptation benchmark designed to shift the STVG evaluation paradigm from static zero-shot testing to rigorous domain adaptation. Targeting five specialized domains (animal, industry, sports, surgery, and public security), AnyGroundBench pairs newly captured, expert-annotated videos with established datasets, unifying them through dense, high-fidelity spatio-temporal annotations. Crucially, the benchmark provides dedicated limited training subsets, enabling systematic evaluation of domain adaptability under limited training data. We benchmark 23 state-of-the-art VLMs in the zero-shot setting and further evaluate five adaptation strategies, spanning training-free and fine-tuning-based approaches, on representative models, assessing their zero-shot generalization and adaptation capacity. Our results show that current VLMs remain far from practical performance in the zero-shot setting, that training-free adaptation produces highly variable effects depending on the model and domain, and that fine-tuning-based adaptation, though more effective, still falls short of real-world requirements, with gains varying markedly across domains. These findings expose fundamental limitations in current VLMs' spatio-temporal reasoning, pointing to concrete directions for future research.
发表机构
- Keio University(庆应大学)
- Keio AI Research Center(庆应人工智能研究中心)
- Keio University School of Medicine(庆应大学医学部)
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。