arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22218cs.CLcs.AI

为什么大语言模型与人类在评估创造力时会趋同和背离

Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity

Pengzhao Lyu, Yeun Joon Kim, Hanlin Xiao, Yingyue Luna Luan

首次发表
浏览论文内容

中文总结 AI 辅助

研究探讨大语言模型与人类评估创造力时趋同和背离的原因,通过三项研究和六个模型,识别其评估标准及影响,发现一致性取决于判断证据和模型标准,强调内在品质时趋同,需上下文信息时背离,选评估模型很关键。

中文摘要 AI 辅助

尽管大语言模型(LLMs)越来越多地被用作创造力评估工具,但其与人类评估的一致性证据仍参差不齐,这引发了它们的判断何时以及为何与人类判断趋同或背离的问题。通过三项研究和六个广泛使用的LLMs,我们识别了LLM创造力评估的潜在标准并研究其下游影响。研究1表明LLMs通常依赖人类创造力评估标准的较窄子集。研究2表明LLM评估与人类评估中度相关,标准更广泛的LLMs能更好地区分人类认为更具创造性和创造性较低的想法。研究3表明LLMs对上下文信息不太敏感。我们的发现有助于解释关于LLM与人类一致性的复杂证据,表明一致性取决于判断所需证据和每个模型应用的标准。

英文摘要

Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from those of humans. Across three studies and six widely used LLMs, we addressed this gap by identifying the standards underlying LLM creativity evaluation and examining their downstream implications. Study 1 showed that LLMs generally relied on a narrower subset of human creativity evaluation standards. Convergence with human standards was strongest in the novelty dimension, whereas divergence was clearest in the contextual dimension, which captures social, market, and reputational information. Moreover, each LLM exhibited distinct, model-specific standards that varied substantially in breadth. These differences in evaluation standards were reflected in actual creativity judgments. Study 2 (N = 1,103 ideas) showed that LLM evaluations were moderately correlated with human evaluations, and individual LLMs with broader standards better distinguished ideas humans judged as more versus less creative. Study 3 (N = 1,195 participants) showed that LLMs were less sensitive to contextual information: such information significantly altered human creativity ratings but left LLM ratings largely unchanged. Together, our findings help explain the mixed evidence on LLM-human alignment, showing that alignment depends on the evidence a judgment demands and the standards each model applies. Selecting an LLM evaluator is therefore a consequential decision: different models, applying different standards, recognise different ideas as creative.

↑