3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
3DRS: MLLMs 需要 3D 意识表示监督以实现场景理解
机构 * Visual AI Lab, The University of Hong Kong(香港大学视觉人工智能实验室) ; Department of Computer Vision Technology (VIS), Baidu Inc.(百度公司计算机视觉技术部)
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
AI总结 3DRS 通过引入预训练 3D 基础模型的监督,提升 MLLM 的 3D 表示能力,从而增强场景理解性能