GeoWorld-VLM: Geometry from World Models for Vision-Language Models
GeoWorld-VLM:从世界模型中获取几何结构用于视觉-语言模型
机构 * Harvard AI and Robotics Lab(哈佛人工智能与机器人实验室) ; Kempner Institute for the Study of Natural and Artificial Intelligence(凯普纳自然与人工智能研究 institute) ; Harvard University(哈佛大学)
专题命中 通用世界模型 :world model(title,abstract);world models(title,abstract);world model(title,abstract);world models(title,abstract)
AI总结 GeoWorld-VLM通过将冻结的摄像机条件视频世界模型的几何结构转移到视觉-语言模型中,提升空间关系推理能力,实验显示在两个不同架构上均提升了约4%的性能。