arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mover360:360°全景图像中的可控物体操作

Mover360: Controllable Object Manipulation in 360° Panoramic Images

Haoyi Zhong, Fang-Lue Zhang, Andrew Chalmers, Taehyun Rhee

arXiv 2608.23238首次发表:更新:

发表机构

Victoria University of Wellington; University of New South Wales; The University of Melbourne(惠灵顿维多利亚大学; 新南威尔士大学; 墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Mover360是针对360°全景图像的可控物体操作框架,以物体平移为核心,支持插入、移除辅助任务,在多域多评估协议下优于相关强基线,代码与基准数据集已公开。

AI 中文摘要

我们提出了Mover360,这是一个针对360°图像的可控物体操作框架。与透视图像不同,等距圆柱投影(ERP)格式的360°图像存在水平环绕、纬度相关失真以及全局场景连续性的特点,这使得现有的透视图像编辑器难以生成符合要求的物体级编辑结果,也让用户难以明确指定编辑需求。为解决这一问题,Mover360以物体平移(在现有全景图中重新定位指定物体)为核心任务,同时支持参考引导的插入和移除作为辅助任务。其界面通过将每个任务编码为固定提示和与ERP对齐的紧凑指令图,统一了点引导、边界框引导和掩码引导的控制方式。在默认点模式下,单次点击即可重新定位物体,模型会利用全景上下文和辅助深度条件推断合理的物体尺寸、支撑关系和光照条件。从结构上看,Mover360是预训练扩散Transformer的轻量级适配模型。为生成配对监督数据,我们构建了UE5数据生成管线,该管线具备表面感知的物体放置和随机光照设置,可生成大规模配对数据以及包含所有三个任务真值的合成与真实全景图的双域基准。在两个测试域和两种评估协议下,Mover360在重建保真度、语义一致性和分布质量方面均优于透视图像编辑、插入和修复任务的强基线方法。代码和我们的基准数据集可在该https链接获取。

英文摘要

We present Mover360, a controllable object manipulation framework for 360° images. Unlike perspective images, 360° images in equirectangular projection (ERP) exhibit horizontal wrap-around, latitude-dependent distortion, and global scene continuity, which makes object-level edits difficult for existing perspective editors to produce and for users to specify. To address this, Mover360 centers on object Translation (relocating a specified object within an existing panorama) while supporting reference-guided Insert and Remove as auxiliary tasks. Its interface unifies point-, bbox-, and mask-guided control by encoding each task into a fixed prompt and a compact, ERP-aligned instruction map. In the default point mode, a single click relocates an object, allowing the model to infer a plausible size, support, and illumination using panoramic context and an auxiliary depth condition. Structurally, Mover360 is a lightweight adaptation of a pretrained diffusion transformer. To generate paired supervision, we construct a UE5 data-generation pipeline with surface-aware object placement and randomized illumination, yielding large-scale paired data and a dual-domain benchmark of synthetic and real panoramas with ground truth for all three tasks. Across both test domains and two evaluation protocols, Mover360 outperforms strong baselines for perspective editing, insertion, and inpainting in reconstruction fidelity, semantic consistency, and distributional quality. Code and our benchmark dataset are available at https://zhonghaoyi.github.io/Mover360/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑