Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial Imagery
Geo3DVQA:基于航拍影像评估视觉-语言模型在三维地理空间推理中的表现
机构 * The University of Tokyo(东京大学) ; RIKEN AIP(日本理化学研究所AIP)
专题命中 视觉空间推理 :reasoning(title,abstract);planning(abstract)
AI总结 Geo3DVQA通过评估视觉-语言模型在三维地理空间推理中的表现,揭示了RGB影像在3D空间分析中的局限性,并展示了领域特定指令微调对模型性能的提升。
Comments Accepted at WACV 2026. 32 pages long including the appendix. Revision details are provided in the supplements