提供维度而非判决:基于评分标准分解的视觉语言美学评委融合
Feed the Panel Dimensions, Not Verdicts: Rubric-Decomposed Fusion of Vision-Language Aesthetic Judges
浏览论文内容
中文总结 AI 辅助
研究发现视觉语言模型整体评判的组合并不优于最佳成员,提出按人工评分标准维度打分并融合,显著提升美学评判准确性。
中文摘要 AI 辅助
视觉语言模型(VLM)被部署为零样本图像美学评判者,且基于薄弱的证据,推荐使用多个模型的组合(panel)来使此类评判者更可靠。在两个人类评分数据集EVA和PARA上,我们发现,无论判决是简单平均还是由学习型组合器融合,整体评判者的组合从未显著优于其最佳成员。组合的价值取决于其输入内容。因此,我们让每个模型依据一份固定的人工编写的评分标准(rubric)的五个维度对每张图像打分,并使用跨模型族的折外(out-of-fold)组合器,将这些维度分数与每个模型的整体判决一同融合。维度分数确实衡量了其标签所声称的内容:在剔除整体人类评分的影响后,在30个模型-属性组合单元中,有28个单元的维度提示比整体提示携带更多特定于属性的信息。融合后,在EVA数据集上,所有十个三模型组合均优于最佳单个VLM(相对于该最佳单个模型,最强三人组合的Spearman rho提升+0.07,预先声明的组合提升+0.10;在二十折划分上平均时,分别提升+0.06和+0.07;相对于组合均值,主要测试在其EVA设计集上提升+0.118),而在PARA数据集上,在Spearman rho下达到持平,在Kendall tau-b下有小幅且不显著的损失,其中单个模型已捕获人类噪声上限的85%。这不是特征数量的伪影:给相同的组合器提供等量的纯整体列(从相同的重复中拆分)并不能重现该结果。这种增益的成本是几百个标签(这些标签不能在数据集间迁移)以及EVA上4.8倍的API调用;我们通过配对自助法和Kendall tau-b报告了该结果,同时报告了一次失败的预注册以及那些表现不佳的配置。
英文摘要
Vision-language models (VLMs) are deployed as zero-shot judges of image aesthetics, and panels of several models are recommended, on thin evidence, as the way to make such judges reliable. On two human-rated datasets, EVA and PARA, we find that a panel of holistic judges never significantly beats its best member, whether the verdicts are averaged or fused by a learned combiner. What a panel is worth depends on what it is fed. We therefore have each model score each image on the five dimensions of a frozen, human-written rubric and fuse those scores, alongside each model's verdict, across model families with an out-of-fold combiner. The dimension scores measure what their labels claim: with the overall human score partialled out, a dimension prompt carries more attribute-specific information than the holistic prompt in 28 of 30 model-attribute cells. Fused, they beat the best single VLM in all ten three-family panels on EVA (against that best single model, +0.07 Spearman rho for the strongest trio and +0.10 for the pre-declared one, and +0.06 and +0.07 when averaged over twenty fold partitions; against the panel mean, the primary test gives +0.118 on its EVA design set), and on PARA they reach parity under Spearman rho and a small, non-significant loss under Kendall tau-b, where one model already captures 85% of the human noise ceiling. It is not a feature-count artefact: giving the same combiner an equal number of pure holistic columns, split from the same repetitions, does not reproduce it. The gain costs a few hundred labels, which do not transfer between datasets, and 4.8x the API calls on EVA; we report it with paired bootstraps and Kendall tau-b, alongside a failed pre-registration and the configurations that lost.
发表机构
- Purdue University Fort Wayne(普渡大学韦恩堡分校)
机构由 AI 辅助整理,请以论文原文为准。