发表机构
Harvard University; Northeastern University(哈佛大学; 东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ARCH-B基准,含354道跨11种建筑表征原型的四选一问题,评估25个多模态模型,发现模型在跨媒介识别上弱于人类且难度相关性低,用于诊断视觉对应与表征迁移。
AI 中文摘要
多模态模型日益能够解读视觉环境,但其在照片、平面图、立面图、剖面图和渲染图之间识别同一建筑的能力仍未得到充分刻画。我们提出了ARCH-B基准,包含354道四选一问题,覆盖11种跨表征原型,基于一个包含390万张建筑图像的建筑关联语料库构建,采用视觉相似干扰项、模型引导的难度筛选和人工验证。我们评估了25个多模态模型,并收集了5830份来自非专家人类参与者的回答。模型准确率介于10.45%至83.90%之间,而人类基线为35.35%。模型在混合表征离群检测和照片匹配上表现相对较好,但在平面图到照片的对应关系上较弱。人类与模型在各原型上的难度仅呈弱相关(斯皮尔曼ρ=0.33)。留出评估证实,筛选期间识别的难度可泛化至策展模型之外。ARCH-B为跨建筑媒介的视觉对应与表征迁移提供了诊断性评估。
英文摘要
Multimodal models increasingly interpret visual environments, but their ability to recognize the same building across photographs, floor plans, elevations, sections, and renderings remains poorly characterized. We introduce ARCH-B, a benchmark of 354 four-choice questions across 11 cross-representational archetypes, constructed from a building-linked corpus of 3.9 million architectural images using visually similar distractors, model-guided difficulty screening, and manual validation. We evaluate 25 multimodal models and collect 5,830 responses from non-expert human participants. Model accuracy ranges from 10.45% to 83.90%, compared with a human baseline of 35.35%. Models perform comparatively well on mixed-representation outlier detection and photograph matching, but remain weaker on floorplan-to-photograph correspondence. Human and model difficulty across archetypes is only weakly correlated (Spearman's (ρ=0.33)). Held-out evaluation confirms that the difficulty identified during screening generalizes beyond the curation models. ARCH-B provides a diagnostic evaluation of visual correspondence and representation transfer across architectural media.