大语言模型智能体能进行发现吗?在机器学习工程任务上评估创造力
Can LLM Agents Discover? Evaluating Creativity on ML Engineering Tasks
浏览论文内容
中文总结 AI 辅助
本研究提出评估LLM智能体创造力的框架,以10个ML工程任务测试AIDE等智能体,发现其P-Creativity随探索转利用下降、H-Creativity高于人类但性能更低,明确其能探索新颖方案却难转化为性能提升。
中文摘要 AI 辅助
近期的AI系统承诺实现自主科学发现,声称能发现算法并产出研究论文,但它们是否具备创造力——即产生既新颖又有用的解决方案的能力——仍是一个悬而未决的问题。我们提出一个框架,以机器学习工程任务为测试平台,评估多轮大语言模型(LLM)研究智能体的创造力,该框架包含三个维度:心理创造力(P-Creativity,指在一次运行中相对于智能体自身先前解决方案的新颖性)、历史创造力(H-Creativity,指相对于人类解决方案语料库的新颖性)以及有用性(任务表现)。我们在来自MLE-Bench的10个Kaggle式机器学习任务上评估了AIDE和AIRA-Dojo这两个智能体框架,开发了一个LLM作为评判者的流程,并验证其与人类创造力判断具有强相关性,为大规模P-Creativity评估提供了可靠的自动化指标。将该流程应用于智能体轨迹后,我们发现:(1)所有智能体在从探索转向利用的过程中,P-Creativity均下降;(2)LLM的H-Creativity高于获得奖牌的人类,但任务表现更低。我们的研究结果表明,当前的智能体能够探索解决方案空间的新颖区域,但缺乏将这种新颖性转化为任务性能提升的能力。
英文摘要
Recent AI systems promise autonomous scientific discovery, claiming to discover algorithms and produce research papers, yet understanding whether they exhibit creativity, the capacity to produce solutions that are both novel and useful, remains an open question. We present a framework for evaluating multi-turn LLM research agents' creativity using ML engineering tasks as a testbed, through three dimensions: P-Creativity (psychological novelty: novel relative to the agent's own prior solutions within a run), H-Creativity (historical novelty: novel relative to the corpus of human solutions), and Usefulness (task performance). Evaluating two agent frameworks, AIDE and AIRA-Dojo, on 10 Kaggle-style machine learning tasks from MLE-Bench, we develop an LLM-as-a-Judge pipeline and verify its strong correlation with human creativity judgments, providing a reliable automated metric for P-Creativity evaluation at scale. Applying this pipeline to agent trajectories, we find: (1) all agents exhibit declining P-Creativity as they transition from exploration to exploitation; (2) LLMs exhibit greater H-Creativity than medal-winning humans, yet achieve lower performance. Our findings reveal that current agents can explore novel regions of the solution space but lack the capacity to convert this novelty into improved task performance.
发表机构
- University of Michigan(密歇根大学)
- Computer Science and Engineering(计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。