发表机构
Sun Yat-sen University; South China University of Technology; Kuaishou Technology; Shenzhen Loop Area Institute(中山大学; 华南理工大学; 快手科技; 深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SpatialDiff通过隐式三维空间建模和全局空间监督,解决了现有图像编辑方法难以处理复杂场景中物体空间运动的问题,提升了空间运动的准确性与保真度。
AI 中文摘要
图像编辑领域的最新进展实现了令人印象深刻的物体操作,但现有方法仍难以处理复杂场景中的空间运动,例如物体跨越不同深度层或部分被遮挡。大多数图像编辑方法仅依赖二维数据集的先验信息,强调平面特征,而缺乏对空间结构的支持。即使是结合了显式位置信息的方法,也无法捕捉真实的三维空间关系,从而限制了复杂场景中物体运动的准确性。本文提出SpatialDiff方法,该方法能有效捕捉三维空间结构,实现复杂场景中精确且一致的物体运动。核心创新有两点:(1)隐式三维空间建模,引入三维先验知识,使模型内部构建对三维空间结构的全面理解;(2)全局空间监督,约束潜在空间特征,使模型能感知编辑操作导致的物体空间位置变化。实验结果表明,该方法显著提升了复杂场景中物体空间运动的准确性和保真度。
英文摘要
Recent advances in image editing allow impressive manipulation of objects, existing methods still struggle to handle spatial movement in complex scenes, such as objects span different depth layers or are partially occluded. Most image editing methods focus solely on prior information from 2D datasets, emphasizing planar features while lacking support for spatial structures. Even approaches that incorporate explicit positional information fail to capture true 3D spatial relationships, thus limiting accurate object movement in complex scenes. In this paper, we present SpatialDiff, a method that effectively captures 3D spatial structures, enabling precise and consistent object movements in complex scenes. Our core innovations are twofold: (1) Implicit 3D Spatial Modeling, which introduces 3D prior knowledge and enables the model to internally build a comprehensive understanding of the three-dimensional spatial structure; and (2) Global Spatial Supervision, which constrains the latent spatial features to enable the model to perceive changes in object spatial positions caused by editing operations. Experimental results demonstrate that our method significantly improves the accuracy and fidelity of spatial movement in complex scenes.
CommentsAccepted by CVPR2026