arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08848cs.CVcs.RO

FIRE3D:一分钟内的前馈交互式3D场景重建

FIRE3D: Feed-forward Interactive 3D Scene Reconstruction Within A Minute

发表机构伊利诺伊大学厄巴纳-香槟分校 · 康奈尔大学
查看机构详情
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

Hongchi Xia, Tianhang Cheng, Wei-Chiu Ma, Shenlong Wang

首次发表
浏览论文内容

中文总结 AI 辅助

FIRE3D提出统一前馈框架,从单张RGB图像或视频在分钟内重建仿真就绪的3D场景,无需测试时优化,速度远超现有方法,并实现物体级完整性。

中文摘要 AI 辅助

我们提出了FIRE3D,一个统一的框架,它接受单个RGB图像或随意拍摄的RGB视频,并在不到一分钟内将其转换为适用于游戏和交互式应用的仿真就绪3D场景资产。FIRE3D的核心是一个前馈、端到端的网络,它从RGB捕获中估计的带姿态的RGB-D观测中预测组合式场景表示,包括每个物体的6自由度姿态、边界框、网格和纹理。通过将场景建模为离散实体的集合,FIRE3D生成模态完整且仿真就绪的环境,其中物体在物理上解耦并准备好进行交互。我们的框架不需要测试时优化,运行速度比先前的交互就绪方法快数个数量级,并提供了超越现有前馈3D方法的物体级完整性。我们在各种数据集上展示了在姿态精度、几何完整性和纹理质量方面具有竞争力或最先进的结果,同时速度要快数个数量级。项目页面:此https URL

英文摘要

We present FIRE3D, a unified framework that transforms a single RGB image or casual video into interactable 3D scene assets for games and interactive applications in under a minute for up to twelve textured instances including preprocessing. At the core of FIRE3D is a feed-forward inference pipeline that predicts a compositional scene representation from posed RGB-D observations estimated from the RGB capture, including the 6-DoF pose, bounding box, mesh, and texture for detected objects. By modeling the scene as a collection of discrete entities, FIRE3D produces amodal object assets that can be independently edited and used in interactive applications. Our framework requires no test-time optimization, runs substantially faster than the compared reconstruction systems, and provides object-level completeness beyond existing feed-forward 3D approaches. We demonstrate leading detection accuracy, strong geometry reconstruction, and competitive rendering quality on the evaluated benchmarks, with substantial runtime gains.

补充信息

↑