发表机构
Qoves Inc.(Qoves公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究对比人类与Claude等4款商业MLLMs的面部吸引力评分,发现MLLMs系统性高估面部吸引力,仅能近似人类评分排序,种族、性别线索判断模式不一致。
AI 中文摘要
多模态大语言模型(MLLMs)的美观度评估正越来越受用户、企业和美学家的青睐,这引发了一个问题:这些AI模型能否准确反映人类对吸引力的判断。在一项预先注册的探索性研究中,我们将2513名人类参与者的吸引力评分与四种广泛使用的商业AI模型(Claude、Gemini、GPT和Grok)的评分进行了比较。结果显示,MLLMs对人脸的评分比人类更积极,且评分范围更窄,在研究进行时,其绝对评分无法复现人类的评分。不过,MLLMs与人类的吸引力判断具有强相关性,能准确追踪人脸的排序。MLLMs可能依据与人类不同的线索判断人脸:年龄是人类和MLLMs判断面部吸引力的共同预测因素,而种族和性别则在不同模型中呈现不一致的模式。AI模型之间的一致性很高,仅Grok例外,它与人类的一致性也最低。我们的研究结果表明,尽管当前现成的商业MLLMs能够近似人类吸引力的排序,但它们系统性地高估了人类面部的美观度。
英文摘要
Beauty assessments from Multimodal Large Language Models (MLLMs) are increasingly popular amongst users, companies, and aestheticians. This raises the question of whether these AI models can accurately reflect human judgments of attractiveness. In a pre- registered exploratory study, we compared the attractiveness ratings of 2,513 human participants to four widely used commercial AI models: Claude, Gemini, GPT, and Grok. Results showed that MLLMs systematically rate faces more favourably and within a narrower range than humans and, at the time of study, do not reproduce human ratings in absolute terms. However, MLLMs exhibit strong correlations with human attractiveness judgments, accurately tracking the rank-ordering of faces. MLLMs may judge faces by different cues than humans; only face age was a predictor of facial attractiveness in both humans and MLLMs, with inconsistent patterns across models for ethnicity and gender. AI models strongly agree with one another, except for Grok, which also showed the lowest agreement with humans. Our findings suggest that while they may be able to approximate rank-orderings of human attractiveness, current off-the-shelf commercial MLLMs systematically overrate the beauty of human faces.