基于CLIP嵌入与SVM的阿联酋住宅建筑多模态文化遗产建筑风格分类
Multimodal Cultural Heritage Architectural Style Classification for Residential Buildings in the UAE Based on CLIP Embeddings and SVM
- University of Sharjah(沙迦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对传统方法依赖图像且缺乏非西方数据的问题,提出基于CLIP的多模态框架融合视觉与文本特征,经UMAP降维、K-Means聚类后训练SVM,在阿联酋住宅建筑八类风格上达到98%准确率,优于现有研究。
AI中文摘要:
文化遗产建筑风格的分析与分类仍然具有挑战性,原因在于建筑视觉图像的复杂性——传统的基于CNN的分类方法高度依赖这些图像而非文本描述——以及缺乏针对非西方地区的特定数据集。本文通过提出一种多模态机器学习框架来解决这一空白,该框架使用OpenAI的CLIP模型对阿联酋住宅建筑进行分析与分类。我们将来自图像的视觉特征与来自专家描述的文本特征整合到一个统一的512维嵌入中,随后使用UMAP进行降维以用于可视化,并使用K-Means进行无监督聚类。通过对K-Means聚类结果进行人工分析得到的聚类标签被用于训练一个SVM分类器,以实现自动化的建筑风格分类。我们的方法在八个识别的风格聚类上达到了98%的分类准确率,高于文献中所有其他研究,证明了结合视觉与文本模态的有效性。总体而言,本文强调了使用多模态AI支持建筑遗产分析的潜力,为探索区域建筑身份提供了可扩展且可解释的工具。
英文摘要:
The analysis and classification of cultural heritage architectural styles remain challenging due to the complexity of visual images of buildings, which are highly relied on in traditional CNN-based classification approaches in comparison to textual descriptions, and the relative lack of non-western region-specific datasets. This paper addresses this gap by proposing a multimodal machine learning framework to analyze and classify Emirati residential architecture using OpenAI's CLIP model. We integrate visual features from images and textual features from expert descriptions into a unified 512-dimensional embedding, followed by dimensionality reduction with UMAP for visualization and unsupervised clustering using K-Means. Cluster labels, which are derived from manual analysis of the K-Means clusters, are used to train an SVM classifier for automated architectural style classification. Our approach achieves a classification accuracy of 98% across eight identified style clusters, higher than every other study in the literature, demonstrating the effectiveness of combining visual and textual modalities. Overall, this paper highlights the potential of using multimodal AI to support architectural heritage analysis, offering scalable and interpretable tools for exploring regional architectural identities.