FIRE3D:一分钟内的前馈交互式3D场景重建
FIRE3D: Feed-forward Interactive 3D Scene Reconstruction Within A Minute
查看机构详情
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
FIRE3D提出统一前馈框架,从单张RGB图像或视频在分钟内重建仿真就绪的3D场景,无需测试时优化,速度远超现有方法,并实现物体级完整性。
中文摘要 AI 辅助
我们提出了FIRE3D,一个统一的框架,它接受单个RGB图像或随意拍摄的RGB视频,并在不到一分钟内将其转换为适用于游戏和交互式应用的仿真就绪3D场景资产。FIRE3D的核心是一个前馈、端到端的网络,它从RGB捕获中估计的带姿态的RGB-D观测中预测组合式场景表示,包括每个物体的6自由度姿态、边界框、网格和纹理。通过将场景建模为离散实体的集合,FIRE3D生成模态完整且仿真就绪的环境,其中物体在物理上解耦并准备好进行交互。我们的框架不需要测试时优化,运行速度比先前的交互就绪方法快数个数量级,并提供了超越现有前馈3D方法的物体级完整性。我们在各种数据集上展示了在姿态精度、几何完整性和纹理质量方面具有竞争力或最先进的结果,同时速度要快数个数量级。项目页面:此https URL
英文摘要
We present FIRE3D, a unified framework that transforms a single RGB image or casual video into interactable 3D scene assets for games and interactive applications in under a minute for up to twelve textured instances including preprocessing. At the core of FIRE3D is a feed-forward inference pipeline that predicts a compositional scene representation from posed RGB-D observations estimated from the RGB capture, including the 6-DoF pose, bounding box, mesh, and texture for detected objects. By modeling the scene as a collection of discrete entities, FIRE3D produces amodal object assets that can be independently edited and used in interactive applications. Our framework requires no test-time optimization, runs substantially faster than the compared reconstruction systems, and provides object-level completeness beyond existing feed-forward 3D approaches. We demonstrate leading detection accuracy, strong geometry reconstruction, and competitive rendering quality on the evaluated benchmarks, with substantial runtime gains.