arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23705cs.CLcs.AIcs.CY

大型语言模型创造力自动评估的局限性

The Limits of Automatic Evaluation of Creativity in Large Language Models

Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi

AI总结:

本研究发现,当前自动评估方法及LLM作为评判者的评估方式,与人类对LLM生成文本创造力的判断存在显著不一致,无法有效捕捉创造力的关键维度。

AI中文摘要:

大型语言模型(Large Language Models, LLMs)在需要创造力的领域生成的文本正越来越多地挑战人类表现,然而对LLM生成内容的创造力进行评估仍是一项重大挑战。本研究调查当前自动评估方法是否能可靠捕捉人类对创造力的判断:从WritingPrompts数据集中收集人类对人类与AI生成的短篇小说在11个创造力维度上的评估,并将这些判断与自动客观指标以及“LLM作为评判者”(LLM-as-a-Judge)的评估进行比较。实验发现自动评估与人类评估之间存在显著不一致,尤其是基于LLM的评判者表现出对AI生成故事的系统性偏好,始终倾向于AI生成文本的风格特征,而非人类创作文本的不可预测性及其他特质。此外,相关性分析显示,广泛使用的自动指标在人类和AI生成的故事中与人类判断的相关性几乎为零,表明这些指标未能捕捉创造力的重要维度。这些发现凸显了当前创意文本自动评估方法的根本性局限,并强调了将创造力的多维主观本质简化为计算指标的难度。

英文摘要:

Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity, yet evaluating creativity in LLM-generated content remains a significant challenge. Here, we investigate whether current automatic evaluation methods can reliably capture human judgments of creativity. We collect human evaluations of human- and AI-generated short stories from the WritingPrompts dataset across 11 dimensions of creativity, and compare these judgments with automated objective metrics and LLM-as-a-Judge evaluations. Our experiments reveal substantial misalignment between automatic evaluations and human assessments. In particular, LLM-based judges exhibit a systematic preference for AI-generated stories, consistently favoring their stylistic characteristics over the unpredictability and other qualities of human-authored texts. Furthermore, correlation analyses show that widely used automatic metrics exhibit near-zero alignment with human judgments across both human- and AI-generated stories, suggesting that they fail to capture important dimensions of creativity. These findings highlight fundamental limitations in current approaches to the automatic evaluation of creative text and underscore the difficulty of reducing the multidimensional and subjective nature of creativity to computational metrics.

↑