发表机构
University of Science and Technology of China; Beijing Institute of Technology; Tsinghua University(中国科学技术大学; 北京理工大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GeoCo-SAVi 通过空间等变解码器、位置对齐和外观移植,实现几何一致的对象表示,显著提升编辑可控性并降低几何误差。
AI 中文摘要
以对象为中心的视频模型用槽表示场景,但显式几何在含义上可能随外观而变化。在不变槽注意力(ISA)中,显式位置和尺度可能与解码的中心和范围不一致;编辑可能导致意外的运动或缩放,替换外观可能改变几何。GeoCo-SAVi 促进几何权威性和语义对齐。其空间等变、对象级解码器使位置和尺度成为有效命令:改变它们会移动或缩放渲染的支持区域。事实位置对齐将位置与解码中心绑定,归一化注意力重叠抑制重复分配。外观移植对齐跨对象的几何语义,使接收者几何控制布局,而捐赠者外观提供形状。时间初始化器跨帧传播校准的槽。在 Obj3D 上,GeoCo-SAVi 匹配 ISA 重建,将 p-中心误差降低超过 80%,并将外观引起的尺寸变化减少超过 50%,同时产生预期的平移和尺度响应。在 250 个 MOVi-C 视频上,它还在重建、实例分组和固定身份几何控制方面优于两个同协议参考。GeoCo-SAVi 将显式几何转化为组合控制,使位置和尺度更可读和可编辑。
英文摘要
Object-centric video models represent scenes with slots, yet exposed geometry can vary in meaning with appearance. In Invariant Slot Attention (ISA), explicit position and scale can disagree with the decoded center and extent; edits can yield unexpected motion or resizing, and replacing appearance can shift geometry. GeoCo-SAVi promotes geometric authority and semantic alignment. Its spatially equivariant, object-wise decoder makes position and scale effective commands: changing them moves or resizes the rendered support. Factual position alignment ties position to the decoded center, and normalized attention overlap discourages duplicate allocation. Appearance transplantation aligns geometry semantics across objects, so recipient geometry governs layout while donor appearance supplies shape. On Obj3D, GeoCo-SAVi matches ISA reconstruction, reduces latent-position-to-decoded-centroid error by nearly 90%, and reduces appearance-induced size variation while producing the expected translation and scale responses. On 250 MOVi-C videos, it improves reconstruction and instance grouping over both same-protocol references, and video editing over STAITUS. GeoCo-SAVi transforms explicit geometry into compositional control, making both position and scale more readable and editable. Project page: https://GeoCo-SAVi.github.io/.
Comments40 pages, 4 figures