发表机构
Tübingen AI Center; University of Tübingen; Meta Reality Labs(蒂宾根人工智能中心; 蒂宾根大学; 元现实实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GenIA是一种无需重训基础模型的测试时输入对齐生成框架,可提升单目、多视图及动态输入下的3D物体位姿预测与重建性能,优于同类方法。
AI 中文摘要
从单目或稀疏多视图观测中重建完整的3D物体资产仍是一项挑战。生成式3D基础模型可补全观测视图之外的物体几何结构,但其预测结果可能无法忠实地复现观测到的几何、外观或位姿。我们提出GenIA,一种测试时输入对齐的生成框架,无需对基础模型进行重新训练,即可将SAM3D的生成先验建立在几何和光度观测的基础上。我们通过从几何中推导平移和缩放并保留学习到的旋转先验来改进物体位姿,在去噪过程中通过可见性偏置注意力、跨观测融合以及可微渲染引导来对齐外观;可选的去噪后细化还可将外观隐变量、轻量解码器适配器和物体放置适配到观测结果。我们的框架还支持外部提供的几何结构,当给定动态物体的时间形状时,它会恢复出共享的、输入对齐的规范外观和稳定的世界空间放置。在合成和真实基准测试中,GenIA在单目、多视图和动态输入下的位姿预测和物体重建性能均有所提升,优于近期基于优化的、逐帧图像转3D以及视频转4D的方法。我们的项目页面可在该https URL获取。
英文摘要
Reconstructing complete 3D object assets from monocular or sparse multi-view observations remains challenging. Generative 3D foundation models can complete object geometry beyond the observed views, but their predictions may not faithfully reproduce the observed geometry, appearance, or pose. We introduce GenIA, a framework for test-time input-aligned generation that grounds SAM3D's generative prior in geometric and photometric observations without retraining the foundation model. We improve object pose by deriving translation and scale from geometry while retaining the learned rotation prior, and align appearance through visibility-biased attention, cross-observation fusion, and differentiable rendering guidance during denoising. An optional post-denoising refinement further adapts the appearance latent, lightweight decoder adapters, and object placement to the observations. Our framework also supports externally supplied geometry; when given temporal shapes of dynamic objects, it recovers a shared, input-aligned canonical appearance and stable world-space placement. Across synthetic and real benchmarks, GenIA improves pose prediction and object reconstruction from monocular, multi-view, and dynamic inputs, outperforming recent optimization-based, per-frame image-to-3D, and video-to-4D methods. Our project page is available at https://facebookresearch.github.io/GenIA.