发表机构
Michigan State University; University of North Carolina at Chapel Hill(密歇根州立大学; 北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RefineAny3D将单目3D检测的深度细化转化为视觉对齐问题,通过VLM实现无需数值预测的深度修正,在多类检测工具上均有性能提升且可泛化。
AI 中文摘要
单目3D目标检测涵盖两种场景:在固定类别词汇内运行的闭集检测器,以及利用深度基础模型获取3D几何信息来定位任意类别的开放词汇检测器。我们发现,当前的深度基础模型尽管具备强大的零样本泛化能力,却缺乏3D检测所需的目标级精度:将最先进的深度基础模型替换为强检测器预测的深度会降低准确率,甚至低于检测器自身的预测结果。我们不追求检测器或深度模型端到端更精准,而是将目标级深度细化作为一项独立任务,提出了RefineAny3D,这是一种无需预测数值的视觉语言模型(VLM),可修正深度。我们的核心见解是:深度误差在图像空间中具有直接的视觉特征:投影到图像上时,位置正确的边界框会紧密包围目标,而距离过远的边界框投影后尺寸过小,距离过近的边界框投影后尺寸过大。因此,深度细化可简化为视觉对齐问题,而非度量回归问题,我们通过以下方式实现:扩展VLM的词汇表,加入动作标记,用分类决策替代数值深度输出;在大规模思维链数据集上监督模型,使每个决策都基于明确的视觉证据。作为单一的后处理步骤应用,RefineAny3D在闭集检测器、开放词汇检测器和3D自动标注工具上均实现了一致的性能提升,且无需重新训练即可泛化到新类别、新场景和新相机。
英文摘要
Monocular 3D object detection spans two regimes: closed-set detectors operating within a fixed category vocabulary, and open-vocabulary detectors that localize arbitrary categories by leveraging depth foundation models for 3D geometry. We find that current depth foundation models, despite their strong zero-shot generalization, lack the object-level precision 3D detection demands: substituting a state-of-the-art depth foundation model for a strong detector's predicted depth degrades accuracy, even falling below the detector's own prediction. Rather than pushing detectors or depth models to be more accurate end-to-end, we treat object-level depth refinement as a stand-alone task and present RefineAny3D, a vision-language model that corrects depth without ever predicting a numerical value. Our key insight is that depth error has a direct visual signature in image space: when projected onto the image, a correctly placed box tightly encloses the object, while a too-far box projects too small and a too-close box projects too large. Depth refinement thus reduces to a visual alignment problem rather than a metric regression problem, which we instantiate by extending the VLM's vocabulary with action tokens that replace numerical depth output with categorical decisions, and by supervising the model on a large-scale chain-of-thought dataset that grounds each decision in explicit visual evidence. Applied as a single post-hoc step, RefineAny3D delivers consistent gains across closed-set detectors, open-vocabulary detectors, and 3D auto-labeling tools, and generalizes to novel categories, scenes, and cameras without retraining.