ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
机构 * University of Illinois Chicago(伊利诺伊大学香槟分校) ; Sony AI(索尼人工智能)
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);LLaVA(abstract);MLLM(abstract)
Comments 8 pages for the main paper