LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
LocateAnything3D: 基于视觉-语言的3D检测与链式视觉推理
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV
AI总结 LocateAnything3D通过将3D检测转化为下一个token预测问题,实现了多目标3D检测,取得了Omni3D基准测试中的最佳成绩。
Comments Tech report. Project page: https://nvlabs.github.io/LocateAnything3D/