发表机构
LSU New Orleans; PinPark, Inc.(路易斯安那州立大学新奥尔良分校; PinPark公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
在 Sim2Real 场景下,研究人员通过实验验证跨视图几何一致性而非单目深度精度决定仅用 RGB 的多相机 3D 跟踪性能,提出几何优先流水线并取得显著优于伪激光雷达方法的结果。
AI 中文摘要
2026 年 AI City Challenge 赛道 1 在合成到真实(Sim2Real)设置下评估大型室内仓库中的多相机 3D 感知;深度信息仅用于训练和验证,因此推理阶段仅使用 RGB。我们采用两条仅用 RGB 的路径作为对一个假设的受控测试:在 Sim2Real 场景下,跨视图几何一致性而非单目深度精度决定性能。第一条是几何优先流水线:YOLO11x 检测、单应性提升至世界坐标系、类别级 3D 尺寸先验、多相机融合、世界坐标跟踪以及离线轨迹段拼接。第二条是估计深度伪激光雷达:将单目深度(D4RT、Metric3D~v2)反向投影到融合点云并输入 3D 检测器(V-DETR),模仿之前使用深度的点云优胜方案。差距显著:几何优先方法达到 13.0 的 3D HOTA(51.6 LocA),而伪激光雷达方法则降至 0.12(9.2 LocA)。我们将性能崩溃归因于单目深度的跨视图不一致性——尺度校正必要但不充分,且域适应微调在预算内无法修复。在几何流水线中,离线拼接是唯一有效的干预措施;SAHI 检测、外观 Re-ID、学习型提升、RT-DETR 集成、测试时增强以及域随机化均未能超越基线检测器。瓶颈互补:检测质量限制几何流水线(DetA),定位一致性限制伪激光雷达(LocA)。我们发布了完整、可复现的仅用 RGB 的流水线及消融实验。
英文摘要
The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) setting; depth is available only for training and validation, so inference is RGB-only. We use two RGB-only routes as a controlled test of one hypothesis: that cross-view geometric consistency, not monocular depth accuracy, governs performance under Sim2Real. The first is a geometry-first pipeline: YOLO11x detection, homography lifting to the world frame, class-level 3D size priors, multi-camera fusion, world-coordinate tracking, and offline tracklet stitching. The second is estimated-depth pseudo-LiDAR: monocular depth (D4RT, Metric3D~v2) back-projected into a fused point cloud and passed to a 3D detector (V-DETR), mirroring prior point-cloud winners that used depth. The gap is decisive: geometry-first reaches 13.0 3D HOTA (51.6 LocA), whereas pseudo-LiDAR collapses to 0.12 (9.2 LocA). We trace the collapse to cross-view inconsistency of monocular depth---scale correction is necessary but not sufficient---which domain-adaptation fine-tuning does not repair within budget. Within the geometry pipeline, offline stitching is the only intervention that helps; SAHI detection, appearance Re-ID, learned lifting, RT-DETR ensembling, test-time augmentation, and domain randomization all fail to beat the baseline detector. The bottlenecks are complementary: detection quality bounds the geometry route (DetA), localization consistency bounds pseudo-LiDAR (LocA). We release a complete, reproducible RGB-only pipeline and ablation.