arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于协同蒸馏发现与双引导鲁棒训练的开放词汇三维目标检测

Open-Vocabulary 3D Object Detection with Co-Distillation Discovery and Dual Guidance Robust Training

Shangbo Yuan, Jie Xu, Xiaofeng Zhu, Na Zhao

arXiv 2608.19973首次发表:更新:

发表机构

School of Computer Science and Engineering, University of Electronic Science and Technology of China; Singapore University of Technology and Design; School of Computer Science and Technology, Hainan University(电子科技大学计算机科学与工程学院; 新加坡科技设计大学; 海南大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出含协同蒸馏发现与双引导训练的开放词汇3D目标检测框架,在SUN RGB-D等数据集上性能优于现有最优方法。

AI 中文摘要

近年来,开放词汇三维目标检测(3D-OVD)因具备在三维场景中检测未见物体的能力而受到越来越多的关注。现有方法通常采用两阶段流程:首先利用基础模型发现新物体,然后基于这些发现的物体训练3D-OVD模型。尽管该方法有效,但此流程在发现阶段常存在定位不准确和分类不匹配的问题,进而限制了模型训练阶段的性能。为解决这些局限,本研究主张同时提升新物体发现的可靠性和模型训练的鲁棒性,并提出了一种创新框架。具体而言,为实现可靠发现,本研究的协同蒸馏策略通过对融合了几何一致性、结构目标性和语义确定性的综合分数应用匈牙利匹配,从而蒸馏出高质量的新物体;为增强模型训练的鲁棒性,本研究进一步提出双引导学习方案,包含针对回归头的场景感知引导不确定性正则化,以及针对分类头的大语言模型(LLM)引导分层对齐,有效缓解了不精确三维边界框和语义模糊性的负面影响。在SUN RGB-D和ScanNetV2数据集上开展的大量实验表明,本方法相较于现有最优方法取得了显著的性能提升,代码可在指定网址获取。

英文摘要

Recently, open-vocabulary 3D object detection (3D-OVD) has gained increasing attention for its ability to detect unseen objects in 3D scenes. Existing approaches typically adopt a two-stage pipeline that first discovers novel objects using foundation models and then trains a 3D-OVD model based on these discovered objects. Although effective, this pipeline often suffers from inaccurate localization and mismatched classification during the discovery stage, which subsequently limits the performance of the model training stage. To address these limitations, we advocate for improving both the reliability of novel object discovery and the robustness of model training, and propose an innovative framework. Specifically, for reliable discovery, our co-distillation strategy distills high-quality novel objects by applying Hungarian matching over a comprehensive score that incorporates geometric consistency, structural objectness, and semantic certainty. To enhance robust model training, we further propose a dual-guidance learning scheme, incorporating a scene-awareness-guided uncertainty regularization for the regression head and an LLM-guided hierarchical alignment for the classification head, effectively mitigating the negative effects of imprecise 3D bounding boxes and semantic ambiguity. Extensive experiments on SUN RGB-D and ScanNetV2 demonstrate that our method achieves significant performance gains over state-of-the-art approaches. Code is available at https://github.com/shangboyuan/Co-3DGT

CommentsAccepted by ECCV26

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑