GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
GeM-VG:迈向通用多图像视觉 grounding 的多模态大语言模型
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 GeM-VG通过引入MG-Data-240K数据集和混合强化微调策略,提升多图像视觉 grounding 和通用多图像理解能力。