物理AI智能空间:面向智能空间多相机三维感知的大规模基准
Physical AI Smart Spaces: A Large-Scale Benchmark for Multi-Camera 3D Perception in Smart Spaces
浏览论文内容
中文总结 AI 辅助
该基准提供大规模多相机三维感知数据,包含280小时视频与近1800相机标注,并提出三维HOTA评估方法,推动智能空间多类别三维框跟踪研究。
中文摘要 AI 辅助
据我们所知,物理AI智能空间(Physical AI Smart Spaces)是首个同时为室内智能空间提供大规模、多类别、多相机三维感知数据的基准。该基准包含由仓库、医院、零售场所等类似场景中近1800台相机拍摄的超过280小时的同步1080p视频,并配有自动标注的多相机身份、二维边界框、三维边界框、相机标定以及可用的深度信息。该基准涵盖了Isaac Sim合成生成、Cosmos Transfer外观增强以及真实世界Sim2Real评估。针对真实世界目标,我们包含两个仓库部署,具有时间同步流、基于VGGT的自动标定,以及一个三维标注界面,该界面将世界坐标系的三维框投影到每个视图以进行跨相机验证。我们描述了数据集范围、标注与标定模式、生成工作流、基准协议以及官方评估系统,该系统标准化了提交格式和排行榜报告。一个核心贡献是三维实例化的高阶跟踪精度(HOTA),将通常基于二维框的跟踪评估扩展到三维位置和三维框。我们进一步报告了来自AI城市挑战赛排行榜的经验基线,展示了方法如何在现实智能空间约束下,从仅限人的三维位置跟踪演进到多类别三维框跟踪。该发布可在以下网址获取:此https网址。
英文摘要
Physical AI Smart Spaces is, to the best of our knowledge, the first benchmark to simultaneously provide large-scale, multi-class, and multi-camera 3D perception data for indoor smart spaces. It contains over 280 hours of synchronized 1080p footage captured by nearly 1,800 cameras in warehouses, hospitals, retail venues, and similar settings, together with automatic annotations for multi-camera identities, 2D bounding boxes, 3D bounding boxes, camera calibration, and depth where available. The benchmark spans Isaac Sim synthetic generation, Cosmos Transfer appearance augmentation, and real-world Sim2Real evaluation. For the real-world target, we include two warehouse deployments with time-synchronized streams, automatic VGGT-based calibration, and a 3D labeling interface that projects world-frame 3D boxes into each view for cross-camera verification. We describe the dataset scope, annotation and calibration schema, generation workflow, benchmark protocols, and official evaluation system, which standardizes submission format, and leaderboard reporting. A central contribution is a 3D instantiation of Higher Order Tracking Accuracy (HOTA), extending the usual 2D box-based tracking evaluation to 3D locations and 3D boxes. We further report empirical baselines from the AI City Challenge leaderboards, showing how methods evolve from person-only 3D location tracking to multi-class 3D box tracking under realistic smart-space constraints. The release is available at https://huggingface.co/datasets/nvidia/PhysicalAI-SmartSpaces.
发表机构
- NVIDIA(英伟达)
- Santa Clara University(圣克拉拉大学)
机构由 AI 辅助整理,请以论文原文为准。