arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型如何评判社交吸引力?来自跨多个LLM与人类的基于理论的人格特质评分的证据

How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans

Hasan Mahmud, Khawaja Abaid Ullah, Mohammad Javad Khojasteh, Jamison Heard, Prabu David

arXiv 2608.09717首次发表:更新:

发表机构

Rochester Institute of Technology(罗切斯特理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文探究LLMs能否评估社交吸引力,经三项研究发现,34个LLM评分稳定且与人类排序一致,但LLMs对吸引力与无吸引力档案的评分偏差更大,且两者均无性别呈现的显著整体效应。

AI 中文摘要

大型语言模型(LLMs)越来越多地被用于执行传统上由人类完成的主观评估,但其作为社交评判者的有效性仍不明确。本文探究LLMs是否能从基于理论的人格特质档案中评估社交吸引力,这些档案由10个心理和关系构念构建,分为三个层级:社交吸引力型、混合型、社交无吸引力型。我们通过两项研究考察LLMs的评分,并在第三项研究中将其与人类判断对比。研究1中,34个LLM对12份档案进行三轮重复评分,尽管部分模型整体评分偏高或偏低,但它们在重复评分中表现出强稳定性、一致的三层级排序,且在档案相对排序上高度一致。研究2使用6对匹配姓名与代词的档案配对,以及一个仅含中性代词的姓名测试,考察对性别呈现的敏感性,两项分析均未发现显著效应。研究3中,198名人类参与者评估了研究2中的6份匹配档案,其评分重现了三层级结构,且排序与LLMs一致;不过,LLMs对吸引力档案的评分比人类更积极,对无吸引力档案的评分更消极,且两组均未显示性别呈现的整体显著效应。

英文摘要

Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social attraction from theory-grounded persona profiles constructed from ten psychological and relational constructs and organized into three tiers: socially attractive, socially mixed, and socially unattractive. We examine LLM ratings in two studies and compare them with human judgments in a third study. In Study 1, 34 LLMs rated 12 profiles across three repeated runs. Although some models tended to give higher or lower ratings overall, they showed strong stability across runs, consistent three-tier ordering, and high agreement in relative profile ordering. Study 2 examined sensitivity to gender presentation using six matched name-and-pronoun profile pairs and a separate pronoun-only test with a gender-neutral name, finding no significant effects in either analysis. In Study 3, 198 human participants evaluated the six matched profiles from Study 2. Their ratings reproduced the three-tier structure and followed a profile ordering consistent with that of the LLMs. However, LLMs rated attractive profiles more positively and unattractive profiles more negatively than humans, while neither group showed a significant overall effect of gender presentation.

Comments9 pages, 2 figures, 2 tables. Includes technical supplement

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑