arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04383cs.CVcs.AI

什么在移动?用于组合式场景控制的局部运动表示

What Moves? Localized Motion Representations for Compositional Scene Control

  • Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)

机构由 AI 辅助整理,请以论文原文为准。

Frank Fundel, Malek Ben Alaya, Thomas Ressler-Antal, Stefan Andreas Baumann, Björn Ommer

AI总结:

针对现有视频表示缺乏局部运动捕捉的问题,提出可提示的局部运动表示,实现对象级运动迁移与多角色视频局部动作分类,提升可控性且优于对比方法。

AI中文摘要:

现实世界的动态具有内在的组合性:多个实体在同一共享场景中同时移动,每个实体都表现出不同的运动模式。然而,大多数现有的视频表示方法对运动进行全局编码,没有明确捕捉单个实体的局部运动。关键在于,运动是相对于全局参考帧定义的,包括相机运动和场景布局。不过,局部嵌入通常是从裁剪后的图像计算得到,或是在编码后通过掩码特征获得,这会丢失解释运动所需的上下文。为解决该问题,我们引入了一种可提示的局部运动表示,能为用户指定的、由空间掩码定义的区域生成持久嵌入。我们的模型不裁剪输入或掩码特征,而是处理完整视频,并直接基于查询区域对运动编码进行条件设置,从而产生时间一致、可按区域寻址的嵌入,既能分离局部动态,又保留了消歧所需的全局上下文。我们展示了对象级运动迁移,实现了动态场景的可控组合。除生成式控制外,我们的嵌入还支持多角色视频中的局部动作分类。在这两项任务中,我们的方法均提升了可控性,且优于通过裁剪或事后掩码实现的全局表示。项目页面:this https URL

英文摘要:

Real-world dynamics are inherently compositional: multiple entities move simultaneously within a shared scene, each exhibiting distinct motion patterns. Yet current motion representation models entangle the dynamics of different entities, without explicitly capturing localized motion for each individually. Crucially, motion is defined relative to a global reference frame, including camera motion and scene layout. However, localized embeddings are often computed from cropped images or obtained by masking features after encoding, discarding the context needed to interpret motion. To address this, we introduce a promptable localized motion representation that produces persistent embeddings for user-specified regions defined by spatial masks. Rather than cropping the input or masking features, our model processes the full video and conditions motion encoding directly on the queried region. This yields temporally consistent, region-addressable embeddings that isolate local dynamics while retaining the global context required for disambiguation. We demonstrate object-level motion transfer, enabling controlled composition of dynamic scenes. Beyond generative control, our embeddings support localized action classification in multi-actor videos. Across both tasks, our approach improves controllability and outperforms global representations localized through cropping or post-hoc masking.

补充信息

↑