arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

数字孪生驱动的Real2Sim2Real:通过配对驾驶场景重建进行模拟器条件生成

Digital Twin-Driven Real2Sim2Real: Simulator-Conditioned Generation via Paired Driving-Scene Reconstruction

Hojun Lim, Hyeongseok Jeon, Donghyun Kim, Soonyoung Jung, Heecheol Yoo

arXiv 2610.08339首次发表:更新:

发表机构

Technical University of Munich; MORAI Inc.; Hyundai Motor Company; Seoul National University(慕尼黑工业大学; MORAI公司; 现代汽车公司; 首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出数字孪生驱动的Real2Sim2Real流水线,通过地理参考数字孪生重建驾驶场景并条件化扩散模型,生成带免费标注的逼真图像,显著替代目标区域真实数据,降低自动驾驶3D感知部署成本。

AI 中文摘要

基于摄像头的自动驾驶3D感知严重依赖大规模标注数据集,将此类系统部署到新的目标区域通常需要数据采集和标注。生成式增强已被提出以降低这一成本,但现有方法面临一个根本性权衡:标签条件方法消耗了它们旨在替代的标注,而模拟器条件方法提供免费标注但缺乏对特定真实环境的视觉基础。本研究探讨了数字孪生驱动的Real2Sim2Real流水线(DT-R2S2R)在多大程度上可以替代目标区域的真实数据。通过在地理参考数字孪生(DT-R2S)中重建记录的驾驶片段,我们以几何对齐的模拟器渲染为条件来训练扩散模型,建立了数字孪生基础的Sim2Real模型(DT-S2R)。因此,DT-S2R在数字孪生覆盖范围内,基于低成本但地理参考的模拟器数据,在重建和新颖模拟器场景中合成逼真的驾驶图像。生成数据的有效性在多种3D检测器上得到验证。特别是DETR3D,在不使用目标图像进行检测器训练的情况下,报告了目标区域真实数据基准(oracle)所获mAP的93.18%。此外,与现有目标外真实数据的简单协同训练优于基准。因此,DT-R2S2R可以大幅降低数字孪生可用区域中手动现场数据采集和标注的成本,为扩展3D感知提供了实用基础。

英文摘要

Camera-based 3D perception for autonomous driving relies heavily on large annotated datasets, and deploying such a system to a new target region typically requires data collection and annotation. Generative augmentation has been proposed to reduce this cost, but existing approaches face a fundamental trade-off: label-conditioned methods consume the very annotations they aim to replace, while simulator-conditioned methods offer free annotations but lack visual grounding to specific real environments. This work investigates the extent to which a digital-twin-driven Real2Sim2Real pipeline (DT-R2S2R) can substitute for target-region real data. By reconstructing recorded driving clips inside a georeferenced digital twin (DT-R2S), we condition a diffusion model on geometrically aligned simulator renderings, establishing a digital twin-grounded Sim2Real model (DT-S2R). As a result, DT-S2R synthesizes photorealistic driving images given low-cost yet georeferenced simulator data across both reconstructed and novel simulator scenes within digital-twin coverage. The efficacy of generated data is verified on diverse 3D detectors. DETR3D, especially, reports 93.18% of mAP obtained by a target-region real-data oracle, without employing target images for detector training. Furthermore, simple co-training with existing out-of-target real data outperforms the oracle. Thus, DT-R2S2R can substantially reduce the cost of manual on-site data collection and annotation in digital twin-available districts, providing a practical foundation for scaling 3D perception.

Comments8 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑