arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过艺术史智能体开展绘画的风格分析

Conducting Stylistic Analysis of Paintings through an Art-History Agent

Marc S. Walton, Astrid Harth

arXiv 2608.29644首次发表:更新:

发表机构

The University of Hong Kong; City University Hong Kong(香港大学; 香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出一种结合视觉Transformer与大型语言模型的艺术史智能体框架,可自动化绘画风格分析,弥合AI概率分类与艺术史风格分析间的方法学鸿沟,推动计算艺术史发展。

AI 中文摘要

将艺术作品归属于某位艺术家,传统上依赖于详细的视觉观察与描述,这在艺术史中被称为风格分析。相比之下,当前该领域使用的人工智能(AI)模型仅能提供无法解释的概率分类。为弥合这一方法学鸿沟,我们提出一个用于自动化绘画风格分析的AI框架,为强化证据收集、发现与验证提供基础。通过在带有元数据的大型绘画语料库上训练视觉Transformer(ViT),我们的系统将艺术史特定数据编码为嵌入向量。这些表示通过稀疏字典学习被分解为训练集中反复出现的一组共享特征。随后,大型语言模型(LLM)通过检索相关艺术作品及其附带的策展人撰写文本,对每个特征进行解读,并将其综合为反映风格属性的描述。最后,自主协调LLM应用推理与行动(ReAct)框架,对这些特征进行加权、测试与优化,形成关于某件艺术作品的连贯描述或艺术作品间的比较。该方法将详细的视觉特征转化为描述性术语,解决了艺术史中的一个关键挑战,从而将图像作为数据的使用与人文学者的语义关切相连接,确立基于视觉的计算艺术史为未来的发展领域。

英文摘要

Attributing an artwork to an artist has traditionally relied on detailed visual observations and descriptions, known as stylistic analysis in art history. By contrast, current artificial intelligence (AI) models used in the field offer only unexplained probabilistic classifications. To bridge this methodological gap, we present an AI framework that automates stylistic analysis of paintings, providing a foundation for enhancing evidence collection, discovery, and verification. By training a vision transformer (ViT) on a large corpus of paintings with metadata, our system encodes this art history-specific data as embeddings. These representations are factorized via sparse dictionary learning into a shared set of features that recur across the training set. A large language model (LLM) then interprets each feature by retrieving associated artworks and their accompanying curator-written texts, and synthesizes them into descriptions that reflect their stylistic attributes. Finally, an autonomous coordinator LLM applies a reasoning-and-action (ReAct) framework to weight, test, and refine these features into cohesive descriptions of an artwork, or comparisons of artworks. This approach converts detailed visual features into descriptive terms, addressing a key challenge in art history. It thus connects the use of images as data with the semantic concerns of humanists, establishing vision-based computational art history as an area for future growth.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑