ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
ViGoR:通过细粒度奖励建模提升大视觉语言模型的视觉 grounding
机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) ; AWS AI(AWS人工智能)
AI总结 ViGoR通过细粒度奖励建模提升大视觉语言模型的视觉 grounding 能力,采用更经济的人类评估和自动化方法,有效提高视觉推理准确性。
Comments Accepted by ECCV 2024