RGBX-R1: Visual Modality Chain-of-Thought Guided Reinforcement Learning for Multimodal Grounding
RGBX-R1: 多视觉模态链式推理引导的多模态接地强化学习
机构 * School of Artificial Intelligence, Tianjin University(天津大学人工智能学院) ; Low-Altitude Intelligence Lab, Xiong’an National Innovation Center(雄安国家创新中心低空智能实验室) ; Xiong’an Guochuang Lantian Technology Co., Ltd.(雄安国创莲田科技有限公司) ; College of Electronic Science and Technology, National University of Defense Technology(国防科技大学电子科学学院)
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
AI总结 RGBX-R1通过视觉模态链式推理引导的强化学习提升多模态接地能力,首次构建RGBX-接地基准并在多个任务中取得显著优势。