发表机构
School of Electronic Information and Communications, Huazhong University of Science and Technology; School of Electrical and Electronic Engineering, Nanyang Technological University(华中科技大学电子信息与通信学院; 南洋理工大学电气与电子工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现有图像美学评估基准的不足,构建含两个互补组件的AesCanvas数据集与基准,评估不同类型MLLMs,发现评论生成与情境敏感判断存在差距,确立文化情境化适配性为美学建模新目标。
AI 中文摘要
多模态大语言模型(MLLMs)的最新进展已将图像美学评估(IAA)从标量评分扩展至可解释的评论与指导。然而现有基准主要评估内在视觉质量或固定领域标准,仍未解决“有吸引力的图像是否适用于特定目的、受众、文化场景或领域惯例”这一问题。我们推出AesCanvas,这一统一套件包含两个互补组件:CritiqueCanvas包含来自54300张图像的519136条指令-响应对,支持摄影、绘画和虚拟图像的长篇多维度评论;ContextCanvas包含301个经专家评审的使用场景,评估现实场景中的情境美学适配性。我们在统一协议下评估了闭源前沿、开放权重通用及美学专用MLLMs。结果显示评论生成与情境敏感判断间存在明显差距:基于参考的词汇和语义指标仅能部分捕捉评论质量,美学专家在选定的评论指标上仍具竞争力,但在ContextCanvas上大幅落后于强大的通用MLLMs。进一步分析表明,美学专长无法可靠迁移至情境适配性,且模型决策可能无法追踪或基于决定性的情境视觉线索。这些发现确立了“文化情境化、基于证据的适配性”作为美学建模的独立目标。
英文摘要
Recent advances in Multimodal Large Language Models (MLLMs) have extended Image Aesthetic Assessment (IAA) beyond scalar scores toward interpretable critique and guidance. Yet existing benchmarks mainly assess intrinsic visual quality or fixed domain criteria, leaving open whether an appealing image is appropriate for a specific purpose, audience, cultural setting, or domain convention. We introduce AesCanvas, a unified suite with two complementary components: CritiqueCanvas with 519,136 instruction-response pairs from 54,300 images supports long-form, multi-dimensional critique across photography, painting, and virtual imagery, whereas ContextCanvas with 301 expert-reviewed use scenarios evaluates contextual aesthetic suitability in realistic use scenarios. Under a unified protocol, we evaluate closed-source frontier, open-weight general, and aesthetic-specific MLLMs. Results reveal a clear separation between critique generation and context-sensitive judgment: reference-based lexical and semantic metrics only partially capture critique quality, while aesthetic specialists remain competitive on selected critique metrics yet substantially lag strong general-purpose MLLMs on ContextCanvas. Further analyses show that aesthetic specialization does not reliably transfer to contextual suitability and that model decisions may fail to track or ground themselves in decisive contextual visual cues. These findings establish culturally situated, evidence-grounded suitability as a distinct objective for aesthetic modeling.
Comments10 pages, 4 figures, 6 tables. Supplementary material included