发表机构
CUHK MMLab; Huawei Research(香港中文大学多媒体实验室; 华为研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出用于评估流式VLMs高动态感知能力的FastBench基准,设计含306个QA对,提出无需训练的基线ProactiveFrame,实验发现现有VLMs在高动态场景下表现受限,为相关研究提供测试平台。
AI 中文摘要
流式视频视觉语言模型(VLMs)支持连续视频理解,但现有基准聚焦于低动态场景。在有限上下文预算下,模型需平衡时间历史、空间分辨率与时间粒度;1-2 FPS的稀疏采样会遗漏快速事件。我们提出FastBench,用于评估真实世界视频流中的高动态感知能力。其基于轨迹的流程包括:从高FPS片段生成问答(QA)、过滤可在2 FPS下回答的问题、使用SAM3与CoTracker3轨迹进行答案验证,以及三轮人工检查。FastBench涵盖8个领域、6种能力,包含306个QA对,具备正向、即时、反向时间范围,且有人工标注的证据区间。我们还提出ProactiveFrame,一种无需训练的基线方法,通过文本令牌调整传入帧率:双层滑动窗口保留近期高FPS观测,同时将旧观测下采样为稀疏历史。实验揭示了显著局限:最强模型Gemini-3.5-Flash仅得50.7分;更密集采样使Qwen3-VL-8B从2 FPS时的32.9分提升至24 FPS时的44.6分,但增益随历史压缩趋于饱和;ProactiveFrame比稀疏均匀采样分别高出5.4和1.5个百分点,但仍远低于神谕引导的聚焦,表明当前VLMs难以仅从流中判断何时需要更精细的时间感知。FastBench为高动态流式视频理解提供了测试平台,代码与数据见此https URL。
英文摘要
Streaming Video Large Language Models (VLMs) enable continuous video understanding, yet existing benchmarks focus on low-dynamic scenarios. Under bounded context budgets, models must balance temporal history, spatial resolution, and temporal granularity; sparse sampling at 1--2 FPS misses fast events. We introduce FastBench to evaluate high-dynamic perception in real-world video streams. Its trajectory-grounded pipeline combines QA generation from high-FPS clips, filtering of questions answerable at 2 FPS, answer verification using SAM3 and CoTracker3 trajectories, and three rounds of human inspection. FastBench contains 306 QA pairs across eight domains, six capabilities, and forward, instant, and backward temporal scopes, with human-annotated evidence intervals. We also present ProactiveFrame, a training-free baseline that adjusts incoming frame rates through text tokens. A dual-tier sliding window retains recent high-FPS observations while downsampling older ones into sparse history. Experiments reveal substantial limitations: the strongest model, Gemini-3.5-Flash, scores only 50.7%. Denser sampling improves Qwen3-VL-8B from 32.9% at 2 FPS to 44.6% at 24 FPS, but gains saturate as history is compressed. ProactiveFrame outperforms sparse uniform sampling by 5.4 and 1.5 percentage points, yet remains well below oracle-guided focusing, showing that current VLMs struggle to determine from the stream alone when finer temporal perception is needed. FastBench provides a testbed for high-dynamic streaming video understanding. Code and data: https://github.com/Ashone3/FastBench.