发表机构
University of Southern California; Stanford University; UC San Diego; Massachusetts Institute of Technology; FortyFive Labs; UC Los Angeles; California Institute of Technology; Harvard University(南加州大学; 斯坦福大学; 加州大学圣迭戈分校; 麻省理工学院; FortyFive实验室; 加州大学洛杉矶分校; 加州理工学院; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该工作提出Neverwhere基准测试套件,含六十多个三维高斯重建场景,用于评估视觉运动控制器,并分析训练数据多样性对性能的必要性。
AI 中文摘要
最先进的视觉运动控制器越来越能够处理复杂的视觉环境,这使得在部署前评估其现实世界性能变得越来越困难。这项工作旨在通过开发一系列超逼真的闭环评估环境——Neverwhere 基准测试套件——来缩小训练与评估之间的差距。该套件包含超过六十个城市室内和室外场景的三维高斯泼溅重建。我们的目标是通过简化基于高斯泼溅的重建在模拟连续测试设置中的创建和集成,鼓励大规模且可复现的机器人评估。我们还通过提供在多个 Neverwhere 场景上训练的策略检查点及其在新场景中评估的性能,强调了仅依赖三维高斯生成数据进行训练的潜在陷阱。我们的分析说明了获取多样化数据以确保性能的必要性。代码和数据可在项目页面获取:此 https URL。
英文摘要
State-of-the-art visual locomotion controllers are increasingly capable at handling complex visual environments, making evaluating their real-world performance before deployment increasingly difficult. This work intends to narrow this train/evaluation gap by developing a collection of hyper-photo-realistic, closed-loop evaluation environments - The Neverwhere Benchmark Suite - comprised of over sixty 3D Gaussian Splatting reconstructions of urban indoor and outdoor scenes. Our goal is to encourage large-scale and reproducible robot evaluation by making it easier to create and integrate Gaussian splats-based reconstructions into simulated continuous testing setups. We also underscore the potential pitfalls of relying exclusively on 3D Gaussian-generated data for training, by providing policy checkpoints trained over multiple Neverwhere scenes and their performance when evaluated in novel scenes. Our analysis illustrates the necessity of sourcing diverse data to ensure performance. Code and data are available on the project page: https://ziyc.github.io/neverwhere-bench/.
Comments9 pages, 14 figures. Accepted to IROS 2026. Project page: https://ziyc.github.io/neverwhere-bench/