arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12179cs.CV

Map-Det3D:面向流输入的多视图3D目标检测的度量前馈3D重建先验

Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs

  • ETH Zürich(苏黎世联邦理工学院)
  • Meta Reality Labs(元宇宙实验室)
  • University of Bonn(波恩大学)

机构由 AI 辅助整理,请以论文原文为准。

Yung-Hsu Yang, Luigi Piccinelli, Samuel Rota Bulò, Sunghwan Hong, Denis Rozumny, Johannes Schönberger, Zuria Bauer, Hermann Blum, Peter Kontschieder, Marc Pollefeys

中文总结 AI 辅助

Map-Det3D是一种在线多视图3D目标检测模型,通过将短时间窗口映射为多视图并利用前馈度量3D重建模型作为几何骨干,直接在度量3D空间预测边界框,实现了稳定的单目视频3D检测,且具有良好的在线性能与鲁棒迁移性。

中文摘要 AI 辅助

度量3D目标检测是具身智能体的核心能力,但多数可靠系统依赖深度传感器,牺牲了成本、功耗与集成简便性,这推动了单目3D检测的发展,它无需额外约束,却面临重大障碍:单张图像的深度,尤其是绝对尺度,约束不足。因此,主流的先检测2D再预测3D属性的模式往往脆弱,因为适度的范围误差就可能主导3D定位,且当相机、运动或环境发生域偏移时,学习到的尺度先验可能失效。为解决此问题,我们提出Map-Det3D,这是一种在线多视图3D目标检测模型,可直接将检测带入从RGB重建的3D空间。我们将短时间窗口映射为多视图,重新利用一个前馈度量3D重建模型作为几何骨干,同时调整其目标感知能力。基于此表示,Map-Det3D直接在度量3D空间中预测边界框,无需广泛使用的2D到3D提升。在不同基准上的实验表明,该设计支持强大的在线性能和无需适应的鲁棒迁移,表明为检测训练重建先验是从单目视频实现稳定度量3D检测的可行途径。代码和模型可在该https URL获取。

英文摘要

Metric 3D object detection is a core capability for embodied agents, yet most reliable systems lean on depth sensors, trading away cost, power, and integration simplicity. This motivates monocular 3D detection, which avoids additional constraints, yet it faces a major obstacle: from a single image, depth, and especially absolute scale, are underconstrained. As a result, the prevailing pattern of detecting in 2D and then predicting 3D attributes is often brittle, since modest range errors can dominate 3D localization, and the learned scale prior can fail when cameras, motion, or environments undergo domain shifts. To address this, we propose Map-Det3D, an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB. We map a short temporal window into multiple views and repurpose a feed-forward metric 3D reconstruction model as our geometric backbone while tuning its object-aware capabilities. Building on this representation, Map-Det3D directly predicts boxes in metric 3D space, without the widely used 2D-to-3D lifting. Experiments across different benchmarks show that this design supports strong online performance and robust transfer without adaptation, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video. Code and models are available at https://royyang0714.github.io/Map-Det3D.

补充信息

相关深度报道

↑