发表机构
Wuhan University; École Polytechnique Fédérale de Lausanne (EPFL); University of North Texas; Australian National University(武汉大学; 洛桑联邦理工学院; 北德克萨斯大学; 澳大利亚国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ArticuTable框架,从单张图像重建具身智能体可交互的3D桌面场景,通过GRAM关节建模和PSGSR场景配准实现部件级关节运动与视图一致布局,并构建ArticuTable-100数据集验证性能。
AI 中文摘要
具身智能体受益于将视觉真实感与物理交互性相结合的3D环境。现有的单图像桌面重建方法能够恢复合理的场景几何结构,但通常将物体表示为单一刚体,将交互限制为整体刚性运动,并排除了可执行的部件级关节运动。同时,恢复与输入视图一致的场景布局仍然具有挑战性,因为单个观察可能允许多种合理的姿态-尺度配置。我们提出了ArticuTable,一个单图像3D桌面重建框架,能够同时恢复可执行的部件级关节运动和与输入视图一致的场景布局。在物体建模方面,我们引入了生成鲁棒关节建模(GRAM),该方法将多模态大语言模型引导的关节拟合与语义状态推理相结合,从不完美的整体代理网格中恢复可靠的关节参数和有效的运动范围,从而将其转换为可执行的关节资产。在场景布局方面,我们提出了渐进式语义-几何场景配准(PSGSR),该方法在互补的度量、平面和输入视图约束下逐步缩小姿态-尺度搜索空间,并通过结构感知的语义对应关系解决方向歧义,从而产生与输入视图一致的场景布局。我们进一步贡献了ArticuTable-100,一个包含100个仿真就绪桌面场景的精选集合。包括用户研究在内的大量评估表明,该方法在视觉保真度、输入视图一致性、关节质量、物理合理性和仿真就绪性方面均表现出强劲性能。
英文摘要
Embodied agents benefit from 3D environments that combine visual fidelity to real-world observations with physical interactivity. Existing single-image tabletop reconstruction methods recover plausible scene geometry but typically represent objects as monolithic rigid bodies, limiting interaction to whole-object rigid motion and precluding executable part-level articulation. Meanwhile, recovering a scene layout consistent with the input view remains challenging because a single observation may admit multiple plausible pose-scale configurations. We present ArticuTable, a single-image 3D tabletop reconstruction framework that recovers both executable part-level articulation and an input-view-consistent scene layout. For object modeling, we introduce generation-robust articulation modeling (GRAM), which combines joint fitting guided by a multimodal large language model with semantic state reasoning to recover reliable joint parameters and valid motion ranges from imperfect monolithic proxy meshes, thereby converting them into executable articulated assets. For scene layout, we introduce progressive semantic-geometric scene registration (PSGSR), which progressively narrows the pose-scale search space under complementary metric, planar, and input-view constraints and resolves orientation ambiguity through structure-aware semantic correspondences, yielding a scene layout consistent with the input view. We further contribute ArticuTable-100, a curated collection of 100 simulation-ready tabletop scenes. Extensive evaluation, including a user study, demonstrates strong performance across visual fidelity, input-view consistency, articulation quality, physical plausibility, and simulation readiness.
Comments25 pages