发表机构
University of Zurich(苏黎世大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对多模态大语言模型在视频定位中训练与部署条件不匹配问题,通过实验表明视觉特征提取是稀疏帧输入瓶颈,提出适配特定层及边界感知采样策略,证明训练策略对稀疏帧视频定位比模型规模更关键。
AI 中文摘要
大规模视频平台每小时处理数百万上传内容,需要能定位视频中违规发生时间和地点的审核系统。处理每个帧在大规模下不可行,系统限于每个视频8到16帧的稀疏输入。然而,最先进的多模态大语言模型在数百帧的密集序列上预训练,导致训练和部署条件不匹配,性能严重下降。本文对训练策略进行系统实证研究以弥合时空视频定位差距。结果表明视觉特征提取是稀疏帧输入下的主要瓶颈。仅适配最后三个ViT层(占总参数4%),能达到68.8%的时间平均交并比,超过零样本8B密集输入模型12.8分。语言模型微调收益可忽略或为负。有时间边界时,边界感知采样策略Hybrid16比均匀采样进一步提高时间平均交并比26分。结论是对于稀疏帧视频定位,训练策略比模型规模更重要,微调后的2B模型始终优于零样本8B模型。
英文摘要
Large-scale video platforms process millions of uploads hourly, requiring moderation systems that can localize when and where policy violations occur within each video. Processing every frame is infeasible at scale, so systems are constrained to sparse inputs of 8 to 16 frames per video. Yet state-of-the-art multimodal large language models (MLLMs) are pretrained on dense sequences of hundreds of frames, creating a fundamental mismatch between training and deployment conditions. This mismatch causes severe performance collapse: the Qwen3-VL 8B model drops from 56.0% to 22.3% temporal mIoU when frames are reduced to 16, a 60.2% relative degradation. We present a systematic empirical study of training strategies to close this gap for spatial-temporal video grounding. Our results suggest that visual feature extraction is the dominant bottleneck under sparse-frame inputs. Adapting only the final three ViT layers, 4% of total parameters, achieves 68.8% temporal mIoU and surpasses a zero-shot 8B model using dense inputs by 12.8 points. Language model fine-tuning, by contrast, offers negligible or negative returns. A boundary-aware sampling strategy, Hybrid16, further improves temporal mIoU by 26 points over uniform sampling when temporal boundaries are available. We conclude that for sparse-frame video grounding, training strategy dominates model scale: a fine-tuned 2B model consistently outperforms a zero-shot 8B model, with or without dense frame access.
Comments15 pages, 5 figures