发表机构
Tsinghua University; Nankai University; University of Electronic Science and Technology of China; Sun Yat-sen University; Nanyang Technological University(清华大学; 南开大学; 电子科技大学; 中山大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在实现引擎原生可编辑3D世界重建。核心方法是引入Lumera管道,基于Lumera-2K数据,用Lumera-Box和Lumera-Light解析相关元素并组装。主要贡献是确立参数化灯光目标,揭示多方面剩余瓶颈,且在检测等方面取得一定分数。
AI 中文摘要
可编辑的3D场景创建需要可检查、移动并导入标准引擎的对象实例和灯光,但现有的单图像方法大多止步于房间规模的几何形状、烘焙/全局光照或文本驱动的生成。我们引入了Lumera(光感知统一引擎原生重建与组装),这是一个用于从单图像进行引擎原生、光感知3D场景解析的基准和参考管道。Lumera-2K由2513个UE5项目构建而成,提供了373万个组件、6300万个对象实例、102600个引擎原生参数化灯光和95100个相机视图。在这些数据上,Lumera-Box和Lumera-Light改编VLM来解析对象框和参数化灯光元组(x,y,z,r,g,b,I),通过逐对象网格重建、HDR环境估计和有界智能细化循环进行组装。在与DetAny3D、SpatialLM、N3D-VLM和WildDet3D的净化框基准测试中,Lumera-Box获得了最强的整体检测、几何、语义和布局分数(合并mAP为0.1141,IoU-B为0.2472,F分数为0.2762),而WildDet3D在锚点召回方面更强。对于灯光,Lumera-Light几乎能恢复所有非空场景(召回率为0.998),但在单个灯光定位方面仍有局限(在0.5米处F1为0.209);匹配灯光的中位位置误差为0.261米,中位ΔE2000为4.59,强度皮尔逊相关系数r = 0.628。这些结果将参数化灯光确立为可测量的可编辑场景目标,并揭示了关系结构、灯光召回/强度和跨引擎泛化方面的剩余瓶颈。
英文摘要
Editable 3D scene creation requires object instances and lights that can be inspected, moved, and imported into standard engines, yet existing single-image methods largely stop at room-scale geometry, baked/global illumination, or text-driven generation. We introduce Lumera (Light-aware Unified Engine-native Reconstruction and Assembly), a benchmark and reference pipeline for engine-native, light-aware 3D scene parsing from a single image. Lumera-2K is built from 2,513 UE5 projects and provides 3.73M components, 63M object instances, 102.6K engine-native parametric lights, and 95.1K camera views. On this data, Lumera-Box and Lumera-Light adapt VLM to parse object boxes and parametric light tuples (x,y,z,r,g,b,I), which are assembled with per-object mesh reconstruction, HDR environment estimation, and a bounded agentic refinement loop. In a sanitized box benchmark against DetAny3D, SpatialLM, N3D-VLM, and WildDet3D, Lumera-Box obtains the strongest overall detection, geometry, semantic, and layout scores (merged mAP 0.1141, IoU-B 0.2472, F-score 0.2762), while WildDet3D remains stronger on anchor recall. For lights, Lumera-Light recovers almost all non-empty scenes (recall 0.998) but remains limited at individual-light localization (F1 0.209 at 0.5 m); matched lights have median position error 0.261 m, median ΔE2000 4.59, and intensity Pearson r=0.628. These results establish parametric lights as a measurable editable-scene target and expose remaining bottlenecks in relation structure, light recall/intensity, and cross-engine generalization.
Comments18 pages, 7 figures, Project Page: https://haidilao0328.github.io/Lumera/