arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30821cs.CVcs.AI

Lucida:用于可组合真实到仿真场景建模的解析、生成与放置

Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

Minghan Qin, Yuang Wang, Xiuyu Yang, Yushi Long, Yujian Zhang, Ruihuan Wang, Kai Ye, Yangang Zhang, Hang Li

首次发表
浏览论文内容

中文总结 AI 辅助

Lucida是一种可组合真实到仿真场景建模方法,通过重新分配流程要求,在三个步骤中利用真实采集的可靠信息,结合GizmoAct策略提升了场景建模相关任务的性能。

中文摘要 AI 辅助

可组合场景建模旨在将真实室内场景恢复为完整、可编辑的对象资产,按观测到的方式排列,为机器人仿真和具身人工智能提供可直接用于仿真的真实环境副本,其中的对象可被单独操控。现有流程将该任务分解为三个步骤:将观测结果解析为实例、为每个实例生成资产、将每个资产放回,但每个步骤都需要杂乱的采集场景极少能提供的输入:准确的实例几何结构、无遮挡视图、与观测结果精确匹配的资产。我们提出Lucida,它保留上述顺序但重新分配了要求,使每个步骤仅使用真实采集能可靠提供的内容,精度在流程末端达到而非在流程起始就要求。Lucida将视频解析为场景图,其节点承载每个实例的多视图证据;根据每个实例的证据生成完整资产;通过GizmoAct放置资产,这是一种视觉语言模型(VLM)策略,将放置视为多轮图形用户界面(GUI)交互,在闭环中操控对象的小部件(gizmo)并自行判断何时对齐完成。在场景级3D对象检测、对象姿态估计和场景重建任务中,Lucida在R2S-Scene上的平均精度均值(mAP)比Boxer提升69%,在CA-1M上将ADD-SB@0.05从57.8%提高到83.4%,将场景F分数从SAM3D的0.794提升至0.924。

英文摘要

Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observed, giving robot simulation and embodied AI a simulation-ready replica of the real environment whose objects can be manipulated individually. Existing pipelines decompose the task into three steps---parse the observations into instances, generate an asset for each, and place each asset back---but every step presumes an input that a cluttered capture rarely provides: accurate instance geometry, unoccluded views, and assets that accurately match the observations. We propose Lucida, which keeps this order but redistributes the requirements, so each step consumes only what a real capture reliably provides and precision is reached at the end of the pipeline rather than demanded at its start. Lucida parses the video into a scene graph whose nodes carry per-instance multi-view evidence, generates a complete asset for each instance from its evidence, and places assets with GizmoAct, a VLM policy that casts placement as multi-turn GUI interaction, manipulating the object's gizmo in a closed loop and deciding itself when alignment is reached. Across scene-level 3D object detection, object pose estimation, and scene reconstruction, Lucida improves mAP over Boxer by 69% on R2S-Scene, raises ADD-SB@0.05 from 57.8% to 83.4% on CA-1M, and increases scene F-Score from 0.794 for SAM3D to 0.924.

发表机构

  • ByteDance Seed(字节跳动Seed)
  • Peking University(北京大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑