ProFocus:通过渐进式视觉聚焦解释艺术图像中的情感体验
ProFocus: Interpreting Affective Experience in Artistic Images with Progressive Visual Focusing
浏览论文内容
中文总结 AI 辅助
本研究针对现有方法无法捕捉艺术情感细微线索的问题,提出ProFocus框架,通过层级艺术评论家与渐进式提示融合模块,在ArtEmis数据集上实现了更优的情感识别与解释性能。
中文摘要 AI 辅助
解释图像触发的情感反应是实现情感智能的核心。与自然图像相比,视觉艺术通过抽象概念和视觉隐喻刻意创作以唤起观者的情感反应,使得情感解释尤为具有挑战性。然而,大多数现有方法依赖通用视觉嵌入(如CLIP),无法捕捉艺术情感背后的细微线索。为解决这一差距,我们提出ProFocus,这是一个通过渐进式视觉聚焦对艺术图像中的情感体验进行建模的新框架。其核心思想是借鉴人类审美欣赏的层级认知理论来建模视觉表示学习。技术上,ProFocus包含两个核心组件:层级艺术评论家(Hierarchical Art Critic,HAC)和渐进式提示融合(Progressive Hint Fusion,PHF)模块。HAC利用多模态大语言模型在三个认知层面(氛围风格、叙事主体和具体细节)生成结构化语言先验,从而将艺术感知转化为连贯的语义引导。基于这些先验,PHF不同于传统跨模态融合,而是将层级提示依次注入视觉特征,实现反映人类感知的渐进式聚焦过程。该设计使模型能够捕捉微妙的情感线索并生成更忠实的解释。在ArtEmis v1.0和v2.0数据集上的大量实验表明,ProFocus在情感识别和情感解释方面均持续优于最先进的方法。项目页面:this https URL。
英文摘要
Interpreting the emotional responses triggered by images is central to achieving emotional intelligence. Compared with natural images, visual art is intentionally created to elicit emotional responses from its viewers through abstract concepts and visual metaphors, making affective interpretation particularly challenging. However, most existing methods rely on general-purpose visual embeddings (e.g., CLIP), failing to capture the nuanced cues underlying artistic emotion. To address this gap, we propose \textbf{ProFocus}, a novel framework that models affective experience in artistic images via progressive visual focusing. The key idea is to model visual representation learning inspired by a hierarchical cognitive theory of human aesthetic appreciation. Technically, ProFocus contains two core components: a Hierarchical Art Critic (HAC) and a Progressive Hint Fusion (PHF) module. HAC leverages multimodal large language models to generate structured linguistic priors at three cognitive levels--atmospheric style, narrative subjects, and concrete details--thereby translating artistic perception into coherent semantic guidance. Building upon these priors, PHF departs from conventional cross-modal fusion by sequentially injecting the hierarchical hints into visual features, enabling a progressive focusing process that mirrors human perception. This design allows the model to capture subtle affective cues and produce more faithful explanations. Extensive experiments on the ArtEmis v1.0 and v2.0 datasets demonstrate that ProFocus consistently outperforms state-of-the-art methods in both emotion recognition and affective explanation. Project page: https://github.com/Zhang-Zhiyan/ProFocus.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- Anhui University(安徽大学)
机构由 AI 辅助整理,请以论文原文为准。