arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32813cs.CVcs.CLcs.LG

USAI-Quant:建筑环境中视觉语言模型的定量推理基准

USAI-Quant: A Quantitative Reasoning Benchmark for Vision-Language Models in Built Environments

Dongdong Wang, Qingqi Song, Yuzhou Chen, Deepak Balakrishnan, Ravi Shankar Srinivasan, Shenhao Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉语言模型在遥感影像定量推理上的不足,构建了首个基于美国335个城市的定量基准USAI-Quant,评估发现现有模型在数值推理上表现不佳。

中文摘要 AI 辅助

大型视觉语言模型(VLMs)已成为城市和空间人工智能的强大范式。然而,当前最先进的大型视觉语言模型在遥感影像的定量推理方面仍然存在困难。现有的基准和算法主要基于定性视觉问答(VQA),对视觉语言模型在建筑环境指标上的定量推理能力提供的见解有限。为弥补这一空白,我们开发了定量城市与空间人工智能基准(USAI-Quant),这是首个旨在通过遥感影像定量评估视觉语言模型在建筑环境指标上推理能力的基准。USAI-Quant 从美国最大的 335 个城市中精选数据,将高分辨率遥感影像与定量建筑环境指标对齐。随后,我们通过将 VQA 应用于三个复杂度级别的数十个建筑环境指标,评估了通用型视觉语言模型和遥感视觉语言模型(RS-VLMs)。我们的结果显示,当前最先进的模型在数值推理任务上始终表现不佳。我们进一步跨模型、问题类型和地理位置进行了深入分析,揭示了性能差异和任务特定挑战的见解。

英文摘要

Large vision-language models (VLMs) have emerged as a powerful paradigm for urban and spatial AI. However, current state-of-the-art large VLMs still struggle with quantitative reasoning on remote sensing imagery. Existing benchmarks and algorithms are predominantly based on qualitative Visual Question Answering (VQA), providing limited insights into the quantitative reasoning capabilities of VLMs for built environment metrics. To address this gap, we develop Quantitative Urban and Spatial AI benchmark (USAI-Quant), the first benchmark designed to quantitatively evaluate VLM's reasoning capabilities on built environment metrics via remote sensing imagery. USAI-Quant is curated from the 335 largest U.S. cities, aligning high-resolution remote sensing images with quantitative built environment metrics. We then evaluate both general-purpose and remote sensing VLMs (RS-VLMs) by applying VQAs to tens of built environment metrics across three complexity levels. Our results reveal that current state-of-the-art models consistently fall short on numeric reasoning tasks. We further conduct in-depth analyses across models, question types, and geographic locations, uncovering insights into performance variability and task-specific challenges.

补充信息

↑