arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FunArt:从生成式3D潜变量解码功能结构与关节

FunArt: Decoding Functional Structure and Articulation from Generative 3D Latents

Dennis Rotondi, Abdelrhman Werby, Kai O. Arras

arXiv 2609.20673首次发表:更新:

发表机构

University of Stuttgart; International Max Planck Research School for Intelligent Systems(斯图加特大学; 马克斯·普朗克智能系统国际研究学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FunArt利用生成式3D潜变量和冻结VAE先验,从静态RGB-D观测构建关节感知功能3D场景图,实现部件分割与关节估计,在Articulate3D上取得最优性能。

AI 中文摘要

为了在人类环境中有效操作,机器人必须识别关节物体,分割其可移动和交互部件,并估计其运动学模型。现有的关节场景表示通常从观察到的交互中恢复运动学,而处理静态扫描的方法往往将关节与功能交互元素分离。我们提出了FunArt,一个从单一静态配置下获取的带姿态RGB-D观测中构建关节感知功能3D场景图的框架。FunArt重建物体实例,将其融合几何直接转换为TRELLIS.2的O-Voxel表示,并利用其冻结的稀疏压缩VAE作为结构先验。一个轻量级基于查询的解码器结合紧凑的物体级潜变量与密集的、表面对齐的特征,联合分割可移动部件和功能交互元素,同时估计运动类型、轴、原点和范围。在Articulate3D数据集上,FunArt在可移动部件分割、关节估计和功能元素分割方面均达到最先进性能,无论有无地面真值物体输入。在端到端设置中,它比最强基线在可移动部件上高出1.5 AP_50点,在关节原点和轴联合约束下高出2.8 AP_50点,在功能元素上高出6.7 AP_50点。这些结果表明,生成式3D潜变量编码了可操作的结构线索,可以在物理交互之前初始化机器人感知和规划。

英文摘要

To operate effectively in human environments, robots must identify articulated objects, segment their movable and interactive parts, and estimate their kinematic models. Existing articulated scene representations typically recover kinematics from observed interactions, while methods operating on static scans often decouple articulation from functional interactive elements. We present FunArt, a framework that constructs articulation-aware functional 3D scene graphs from posed RGB-D observations captured in a single static configuration. FunArt reconstructs object instances, converts their fused geometry directly into the O-Voxel representation of TRELLIS.2, and exploits its frozen, sparse-compression VAE as a structural prior. A lightweight query-based decoder combines compact object-level latents with dense, surface-aligned features to jointly segment movable parts and functional interactive elements while estimating motion type, axis, origin, and range. On the Articulate3D dataset, FunArt achieves state-of-the-art performance across movable-part segmentation, articulation estimation, and functional-element segmentation, both with and without ground-truth object input. In the end-to-end setting, it outperforms the strongest baselines by 1.5 AP_{50} points for movable parts, 2.8 AP_{50} points under joint origin-and-axis constraints, and 6.7 AP_{50} points for functional elements. These results demonstrate that generative 3D latents encode actionable structural cues that can initialize robotic perception and planning before physical interaction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑