arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.12684cs.CVcs.AIcs.HC

视觉审美基准:前沿模型能评判美吗?

Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?

  • Bake AI
  • University of Washington(华盛顿大学)
  • University of California, Santa Barbara(加州大学圣巴巴拉分校)
  • Stanford University(斯坦福大学)
  • University of Notre Dame(诺丁汉大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • MIT-IBM Watson AI Lab(麻省理工-IBM沃森人工智能实验室)
  • Western Washington University(西雅图华盛顿大学)
  • King Abdulaziz City for Science and Technology(国王阿卜杜勒阿齐兹科技城)

机构由 AI 辅助整理,请以论文原文为准。

Yichen Feng, Yuetai Li, Chunjiang Liu, Yuanyuan Chen, Fengqing Jiang, Yue Huang, Hang Hua, Zhengqing Yuan, Kaiyuan Zheng, Luyao Niu, Bhaskar Ramasubramanian, Ba… 展开作者

Yichen Feng, Yuetai Li, Chunjiang Liu, Yuanyuan Chen, Fengqing Jiang, Yue Huang, Hang Hua, Zhengqing Yuan, Kaiyuan Zheng, Luyao Niu, Bhaskar Ramasubramanian, Basel Alomair, Xiangliang Zhang, Misha Sra, Zichen Chen, Radha Poovendran, Zhangchen Xu

更新

AI总结:

本文提出视觉审美基准(VAB),通过比较选择评估审美,发现前沿模型在判断最佳和最差图像时表现逊于人类专家,表明需进一步改进。

AI中文摘要:

多模态大语言模型(MLLMs)如今被广泛用于视觉理解、生成和整理。大量应用需要显式审美判断,但现有方法多将判断简化为单张图像的标量评分。本文通过八名专家标注者的控制研究发现,评分派生的排名与直接比较不一致,而直接排名在最佳和最差图像标签上表现出更高的标注者一致性。受此启发,我们引入视觉审美基准(VAB),将审美评估转化为具有匹配主题的候选集中的比较选择。VAB包含400个任务和1195张图像,涵盖细艺术、摄影和插画,标签由每个任务10名独立专家的共识得出。评估20个前沿MLLMs和6个专用视觉质量奖励模型,发现最强系统在三个随机候选顺序排列中仅在26.5%的任务中正确识别最佳和最差图像,远低于人类专家的68.9%。在2000个专家示例上微调35B参数模型使其准确性接近397B参数开放权重模型,表明VAB中的比较信号具有可转移性。这些结果揭示了当前多模态模型与专家审美判断之间的明显且可测量的差距,VAB提供了首个基于集合的、专家导向的测试床,用于追踪和缩小这一差距。

英文摘要:

Multimodal large language models (MLLMs) are now routinely deployed for visual understanding, generation, and curation. A substantial fraction of these applications require an explicit aesthetic judgment. Most existing solutions reduce this judgment to predicting a scalar score for a single image. We first ask whether such scores faithfully capture comparative preference: in a controlled study with eight expert annotators, score-derived rankings align poorly with the same annotators' direct comparisons, while direct ranking yields substantially higher inter-annotator agreement on best- and worst-image labels. Motivated by this finding, we introduce the Visual Aesthetic Benchmark (VAB), which casts aesthetic evaluation as comparative selection over candidate sets with matched subject matter. VAB contains 400 tasks and 1,195 images across fine art, photography, and illustration, with labels derived from the consensus of 10 independent expert judges per task. Evaluating 20 frontier MLLMs and six dedicated visual-quality reward models, we find that the strongest system identifies both the best and the worst image correctly across three random permutations of the candidate order in only 26.5% of tasks, far below the 68.9% achieved by human experts. Fine-tuning a 35B-parameter model on 2,000 expert examples brings its accuracy close to that of a 397B-parameter open-weight model, suggesting that the comparative signal in VAB is transferable. Together, these results expose a clear and measurable gap between current multimodal models and expert aesthetic judgment, and VAB provides the first set-based, expert-grounded testbed on which that gap can be tracked and closed.

补充信息

↑