arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Object-Uni:面向以物体为中心的空间理解与可控生成的统一模型

Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

Mining Tan, Yinuo Wang, Ziqi Zhou, Weize Quan, Sifei Li, Jingdong Chen, DanDan Zheng, Libin Wang, Weiming Dong

arXiv 2608.22757首次发表:更新:

发表机构

School of Artificial Intelligence, University of Chinese Academy of Sciences; MAIS, Institute of Automation, Chinese Academy of Sciences; School of Vehicle and Mobility, Tsinghua University; Ant Group(中国科学院大学人工智能学院; 中国科学院自动化研究所MAIS; 清华大学车辆与 mobility学院; 蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Object-Uni统一模型,通过视角方向抽象与UniSpatial-80K基准,实现以物体为中心的空间理解与可控生成,提升位姿理解及可控生成能力。

AI 中文摘要

视觉理解与生成的统一模型已取得快速进展,但仍缺乏对物体实例空间状态的理解与操控能力。现有模型可通过自然语言描述物体,却难以精确表示连续的物体位姿,也难以在目标视角下生成几何一致的图像。为缓解该问题,我们提出Object-Uni,一种面向以物体为中心的空间理解与可控生成的统一模型。具体而言,我们将以物体为中心的空间智能形式化为连接位姿感知、空间推理、位姿条件生成及以物体为中心的新视角合成的统一问题。我们将物体位姿视为理解与生成共享的显式几何变量,而非仅作为预测标签或控制信号。为使位姿可被多模态大语言模型使用,我们提出一种基于视角的方向抽象方法,将方向映射为结构化的视角描述,同时保留连续几何监督。我们进一步构建了以物体为中心的空间基准UniSpatial-80K,并训练了一个带有以物体token为基础的位姿锚点的统一模型,以将每个实例与其位姿状态关联。实验表明,我们的模型提升了物体级位姿理解与位姿可控生成能力,推动统一模型从描述物体向操控空间状态迈进。

英文摘要

Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects in natural language, but they struggle to precisely represent continuous object poses and generate geometrically consistent images under target viewpoints. To mitigate this, we propose \emph{Object-Uni}, a unified model for object-centric spatial understanding and controllable generation. Specifically, we formulate object-centric spatial intelligence as a unified problem connecting pose perception, spatial reasoning, pose-conditioned generation, and object-centric novel view synthesis. We treat object pose as an explicit geometric variable shared by understanding and generation, rather than merely a prediction label or control signal. To make pose usable by multimodal large language models, we propose a viewpoint-based orientation abstraction that maps orientation into structured viewpoint descriptions while preserving continuous geometric supervision. We further construct an object-centric spatial benchmark (UniSpatial-80K) and train a unified model with an object-token-grounded pose anchor to associate each instance with its pose state. Experiments show that our model improves object-level pose understanding and pose-controllable generation, moving unified models from describing objects toward manipulating spatial states.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑