AI 中文总结
研究探讨视觉语言模型从街景图像推断建筑类型的潜力,将其预测与人类专家比较,评估多种VLM,发现思维链提示性能更稳定,通过分析关键词概率研究推理过程,表明VLM能逼近专家能力,为城市分析提供支持。
AI 中文摘要
本研究探讨视觉语言模型(VLM)从谷歌街景(GSV)图像推断建筑类型(建造、当前用途和层数)的潜力。将VLM生成的预测与人类专家(土木工程师和建筑师)的推断进行比较,作为手动标注的地面真值数据来源。评估了几种先进的VLM,发现思维链提示能提供更稳定的模型性能。通过检查关键词在人工智能解释中出现的概率,研究了VLM建筑类型预测背后的推理。结果表明,人工智能倾向于关注视觉指标,而人类专家除视觉线索外,更强调更广泛的上下文线索和领域知识。总体而言,VLM在建筑类型分类上能大规模逼近专家能力,平均准确率约70%。该研究证明了VLM在城市环境中模式识别和对象识别任务中的人工智能自动化潜力,有助于探索人工智能视觉预测的效率和可扩展性,并为支持城市分析和预测自动化过程的推理过程提供见解。
英文摘要
This research investigates the potential of Vision-Language Models (VLMs) to infer building typologies: Construction, Current Use, and Storeys from Google Street View (GSV) images. Predictions generated by VLMs are compared with inference by human experts (civil engineers and architects) as a source of manually labelled ground-truth data. We evaluate several state-of-the-art VLMs, including GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash. By applying different scaling strategies and prompting techniques, we found that Chain-of-Thought prompts provide an overall more stable model performance. We also investigate the reasoning behind VLMs' building-typology predictions by examining the probabilities of keywords appearing in AI explanations. This enabled us to analyse patterns in these reasonings and identify key themes driving both agreements and disagreements between VLM and expert labels. We find that AI tends to focus on visual indicators, whereas human experts place greater emphasis on broader contextual cues and domain knowledge, in addition to visual cues. Overall, VLM can approximate experts' capability in building-typology classification at scale, with an average accuracy of approximately 70%. The study demonstrates the VLM's potential for AI automation in tasks that require pattern recognition and object identification in an urban context. AI have the potential to serve as complementary and collaborative tools for urban analysis, leveraging their strengths in understanding visual patterns. This study contributes to the exploration of the efficiency and scalability of AI visual prediction and provides insights into the reasoning processes that could support automation processes in urban analysis and prediction.