发表机构
NVIDIA; Simon Fraser University(英伟达; 西蒙菲莎大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GATOR是一种生成式智能体框架,可从一张或多张随手拍摄的图像中恢复带纹理的物体资产及其场景相对位姿,通过局部模态混合器、文本引导语义条件和多模态智能体的编辑-渲染-审查循环,在多场景下实现了高保真重建,兼具效率与模拟就绪性。
AI 中文摘要
从随手拍摄的图像重建完整、与场景对齐的三维物体,需要整合稀疏且不确定的观测结果,并推断被遮挡物隐藏的表面。我们提出GATOR,这是一个生成式智能体框架,可从一张或多张图像中恢复带纹理的物体资产及其相对于场景的位姿。我们的局部模态混合器在跨视图推理前,将补丁对齐的RGB、目标掩码和点图特征进行耦合,在保留场景上下文的同时,区分目标与其周围环境。文本引导的语义条件通过针对结构、几何和外观生成的特定阶段适配器,用类别名称和物体描述补充这些空间线索。生成的资产初始化一个多模态智能体,为通过观测引导的编辑-渲染-审查循环进行针对性的结构和纹理细化,提供实例特定的几何和位姿。在合成物体、杂乱桌面和室内场景上,GATOR实现了强大的几何和外观保真度,同时从稀疏观测中恢复场景相对位姿。时间预算比较和场景级模拟进一步证明了其重建效率和模拟就绪性。项目页面:this https URL
英文摘要
Reconstructing complete, scene-aligned 3D objects from casual images requires integrating sparse, uncertain observations and inferring surfaces hidden by occlusions. We present GATOR, a generative and agentic framework that recovers textured object assets and their scene-relative pose from one or more images. Our local modality mixer couples patch-aligned RGB, target-mask, and pointmap features before cross-view reasoning, preserving scene context while distinguishing the target from its surroundings. Text-guided semantic conditioning complements these spatial cues with category names and object captions through stage-specific adapters for structure, geometry, and appearance generation. The generated asset initializes a multimodal agent, providing instance-specific geometry and pose for targeted structural and texture refinement through an observation-guided edit-render-review loop. Across synthetic objects, cluttered tabletops, and indoor scenes, GATOR achieves strong geometric and appearance fidelity while recovering scene-relative pose from sparse observations. Time-budget comparisons and scene-level simulation further demonstrate the reconstruction efficiency and simulation readiness. Project page: https://research.nvidia.com/labs/lpr/gator/