arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

创造力的秘诀:大语言模型中的迭代生成与评估

Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models

Rens Anderson, Tessa Verhoef, Amirhossein Zohrehvand

arXiv 2608.07243首次发表:更新:

发表机构

Leiden University(莱顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究以2024年Pillsbury烘焙大赛食谱生成为任务,探究迭代生成对LLM创造力的影响,发现迭代生成-选择可产出与人类基准相当的食谱,且评估器设计是主观创造性搜索的关键变量。

AI 中文摘要

生成模型通常通过单一产物进行评估,而人类创造力往往源于迭代生成、评价与完善。本试点研究通过将FunSearch适配至2024年Pillsbury烘焙大赛的食谱生成任务,并采用基于TTCT的LLM评估方法对照人类基准评估输出,探究迭代搜索是否能提升LLM的创造力。在两项实验中,我们测试了迭代次数、生成器温度以及循环内选择评分器模型的规模。结果显示,迭代生成-选择可产生创造力评分与人类基准相当的食谱,但仅增加迭代次数无法提升创造力;循环内评估器最为关键,较小的选择评分器在多数TTCT维度上产生显著更高的评分,而温度除原创性外影响有限。这些发现表明,评估器设计是主观创造性搜索中的一阶设计变量。

英文摘要

Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM creativity by adapting FunSearch to recipe generation for the 2024 Pillsbury Bake-Off and evaluating outputs against human benchmarks using TTCT-based LLM evaluation. Across two experiments, we test iteration count, generator temperature, and in-loop selection-scorer model size. Results show that iterative generation-selection can produce recipes with creativity scores comparable to human benchmarks, but additional iterations alone do not improve creativity. The in-loop evaluator matters most: a smaller selection scorer yields significantly higher scores across most TTCT dimensions, while temperature has limited effects except for originality. These findings suggest that evaluator design is a first-order design variable in subjective creative search.

Comments7 pages, 3 figures, 1 table. Short paper accepted at ICCC'26

Journal refProceedings of the 17th International Conference on Computational Creativity (ICCC'26), Coimbra, Portugal, June 29-July 3, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑