LLaVA$^3$: Representing 3D Scenes like a Cubist Painter to Boost 3D Scene Understanding of VLMs
LLaVA$^3$:像立体主义画家一样表示3D场景以提升VLM的3D场景理解
专题命中 空间理解 :3D reconstruction(abstract);分类 cs.CV
AI总结 LLaVA$^3$通过立体主义方法提升VLM对3D场景的理解能力,利用多视角2D图像无需微调实现更优的3D场景理解。
Comments Accepted at AAAI'26