发表机构
TU Munich; TU Darmstadt; ETH Zurich; University of Cambridge(慕尼黑工业大学; 达姆施塔特工业大学; 苏黎世联邦理工学院; 剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对单目时序3D检测中基于查询的模型泛化能力差的问题,提出首个领域通用单目时序3D检测方法MAGneT-3D,采用DRAG与TRIM策略,在跨数据集基准上零样本域偏移下NDS提升至18.6%。
AI 中文摘要
单目时序3D检测旨在给定单目视频的情况下检测3D空间中的物体。基于查询的3D检测器将检测与跨视图关联统一起来,但其可学习查询会适配训练数据的空间分布(如视野)。我们表明,当这些模型应用于单目视频时,该问题尤为严重,阻碍了对未见数据集和环境的泛化。为解决此局限,我们提出MAGneT-3D,这是首个用于领域通用单目时序3D物体检测的方法。我们不依赖静态可学习查询,而是提出领域鲁棒锚点生成器(DRAG)方法,在推理时自适应生成3D提议。为进一步实现领域泛化,我们提出时序优化与身份合并(TRIM)策略,减少对特定3D提议的依赖。为开展全面的领域泛化评估,我们建立了涵盖nuScenes、Waymo、Lyft和ONCE的跨数据集基准。在零样本域偏移下,MAGneT-3D优于所有基线,将NDS从12.1%提升至18.6%,同时也提高了域内准确率。
英文摘要
Monocular temporal 3D detection aims to detect objects in 3D, given a monocular video. Query-based 3D detectors unify detection and cross-view association, but their learnable queries fit the spatial distribution of the training data (e.g., field-of-view). We show that this issue is especially severe when these models are applied to monocular video, hindering generalization to unseen datasets and environments. To address this limitation, we introduce MAGneT-3D, the first method for domain-generalized monocular temporal 3D object detection. Instead of relying on static learnable queries, we propose a Domain-Robust Anchor Generator (DRAG) approach that adaptively derives 3D proposals during inference. To further enable domain generalization, we propose a Temporal Refinement and Identity Merging (TRIM) strategy, reducing dependence on specific 3D proposals. To enable comprehensive domain-generalization evaluation, we establish a cross-dataset benchmark spanning nuScenes, Waymo, Lyft, and ONCE. Under zero-shot domain shifts, MAGneT-3D outperforms all baselines, improving NDS from 12.1% to 18.6% while also increasing in-domain accuracy.
CommentsTo appear at ECCVW 2026 (DriveX workshop; Oral paper). Johannes Meier and Mohamed Kotb - both authors contributed equally. Project page: https://mo-sameh.github.io/MAGneT-3D-Project-Page/