发表机构
Cornell University; Meta(康奈尔大学; Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ViGeo框架,通过时空画布补全将视觉类比扩展到视频领域,实现无需训练的多样化视频任务统一,并识别和解决了任务内化捷径问题。
AI 中文摘要
将视频模型适应于新任务通常需要专门的数据整理和微调。虽然视觉类比通过上下文指定任务提供了一种无需训练的替代方案,但它仍局限于图像领域。为了探索基于类比的方法能否统一多样化的视频任务并泛化到分布外场景,我们引入了ViGeo,这是一个通过时空画布补全将视觉上下文学习扩展到视频领域的框架。在具有严格训练-测试划分的多样化任务分类法上进行评估,ViGeo能够泛化到未见过的视频操作和零样本模态(例如,事件相机)。最后,我们识别了任务内化现象,即与预训练任务关联的查询格式覆盖了演示,并表明这种捷径可以通过少量与任务无关的数据来消除,这凸显了将提示格式与任务身份去相关化的必要性。
英文摘要
Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion. Evaluated on a diverse task taxonomy with a strict train-test split, ViGeo generalizes to unseen video manipulations and zero-shot modalities (e.g., event cameras). Finally, we identify task internalization, where a query format associated with a pretrained task overrides the demonstration, and show that this shortcut can be removed with a small amount of task-unrelated data, highlighting the need to decorrelate prompt format from task identity.
CommentsProject page: https://iandrover.github.io/video_analogy/