大型语言模型(LLM)正变得同样具有创造力吗?来自三年模型的证据
Are LLMs becoming similarly creative? Evidence from three years of models
浏览论文内容
中文总结 AI 辅助
本文通过分析三年间的LLM在Infinity-Chat100和替代用途任务上的表现,发现LLM创造性输出多样性随时间显著下降,存在同质化趋势,需关注其对人机协同创作中人类能动性的影响。
中文摘要 AI 辅助
许多基准测试追踪大型语言模型(LLM)在具有可验证答案的任务上的性能,但人们对LLM在开放式任务上的性能演变知之甚少,在这类任务中,创造力、原创性和多样性可能与质量同样重要。随着LLM越来越多地支持人类构思和创造性工作,了解LLM在开放式任务上的性能趋势至关重要。本文对跨越三年模型发布的LLM创造性输出进行了初步分析,研究了模型对Infinity-Chat100(真实世界的开放式用户查询集合)和替代用途任务(Alternate Uses Task,一种已确立的心理测量学创造力评估方法)的响应。利用句子嵌入相似度,我们研究了LLM对这些提示的响应趋势。我们的发现显示,随着时间推移,模型输出多样性出现统计上显著的下降,表明不同模型的LLM输出在创造性内容上可能正在趋同。如果这一趋势持续下去,由LLM驱动的同质化可能会逐步削弱人类在人机协同创造性工作中的能动性,需要审慎考量LLM在人类创造性过程中的角色。
英文摘要
Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality. As LLMs increasingly support human ideation and creative work, understanding trends in LLM performance on open-ended tasks is critical. This paper presents a preliminary analysis of LLM creative outputs spanning three years of model releases, examining model responses to Infinity-Chat100, a real-world collection of open-ended user queries, and the Alternate Uses Task, an established psychometric creativity assessment. Using sentence-embedding similarity, we examine trends in LLM responses to these prompts. Our findings show a statistically significant decrease in model output diversity over time, suggesting that LLM outputs may be converging in creative substance across models. If this trend persists, LLM-driven homogenization may progressively diminish human agency in human-AI co-creative work, demanding careful consideration of LLMs' role in the human creative process.
发表机构
- Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。