基于几何知识蒸馏的时序锚定组合式相机运动理解
Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation
浏览论文内容
中文总结 AI 辅助
本研究提出CamDistill方法,通过几何知识蒸馏实现时序锚定的组合式相机运动理解,推出含4229个片段的CamChoreo基准,提升了相机运动识别的细粒度与效率。
中文摘要 AI 辅助
理解相机运动是视频感知的基础,应用于空间智能和可控视频生成领域。多模态大语言模型(MLLMs)为该任务提供了自然接口,但现有工作通常为整个视频片段分配一个或多个标签。这种片段级识别忽略了真实相机运动的两个关键特性:同一镜头内运动可能发生变化,且多种运动可同时发生。因此,我们将相机运动理解定义为时序锚定的组合式识别,要求模型定位运动一致的时间区间,并识别每个区间内的所有活动运动。我们推出CamChoreo基准,包含4229个真实单镜头视频片段,带有专家标注的时间片段。其标注使用由20个方向感知标签组成的精简词汇表,近一半片段包含复合相机运动,即多个运动基元同时活动。当前MLLMs难以识别这种细粒度组合式运动,因其视觉编码器侧重语义内容而非相机运动依赖的几何证据。直接注入冻结3D基础模型的特征可解决该问题,但需对每个输入运行昂贵的几何模型,我们将此基线方法称为CamInject。我们提出CamDistill,在训练期间将相同几何知识蒸馏到轻量相机令牌中,推理阶段移除3D模型。CamDistill达到了直接特征注入的精度,且推理阶段无需运行3D教师模型。CamChoreo与CamDistill共同将相机运动理解从片段级标签任务推进至时序锚定的组合式识别。项目页面:this https URL。
英文摘要
Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation. Multimodal large language models (MLLMs) provide a natural interface for this task, but existing work typically assigns one or more labels to an entire clip. Such clip-level recognition overlooks two defining properties of real camera motion: it can change within a shot, and multiple movements can occur simultaneously. We therefore formulate camera-motion understanding as temporally grounded, compositional recognition, which requires a model to localize motion-consistent intervals and identify every movement active within each interval. We introduce CamChoreo, a benchmark of 4,229 real single-shot clips with expert-annotated temporal segments. Its annotations use a compact vocabulary of 20 direction-aware labels, and nearly half of the segments contain compound camera motion, with multiple movement primitives active simultaneously. Recognizing such fine-grained, compositional motion is hard for current MLLMs, whose visual encoders emphasize semantic content rather than the geometric evidence on which camera motion depends. Directly injecting features from a frozen 3D foundation model addresses this gap, but requires running the expensive geometry model on every input; we refer to this baseline as CamInject. We instead propose CamDistill, which distills the same geometric knowledge into lightweight camera tokens during training and removes the 3D model at inference. CamDistill matches the accuracy of direct feature injection without running the 3D teacher at inference. Together, CamChoreo and CamDistill advance camera-motion understanding from clip-level labeling to temporally grounded, compositional recognition. Project page: https://ddz16.github.io/cammotion.github.io/.
发表机构
- The Hong Kong University of Science and Technology(香港科技大学)
- Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。