当AI设计AI:创新还是模仿?
When AI Designs AI: Innovation or Imitation?
浏览论文内容
中文总结 AI 辅助
该研究对比LLM智能体与人类设计AI方法的性能和算法差异,发现智能体偶达SOTA但无法泛化,96.8%的设计属人类推导空间,多为算法选择的复用重组。
中文摘要 AI 辅助
大型语言模型(LLM)智能体的最新进展使其越来越有能力为复杂AI任务设计方法,这引发了两个关于智能体设计方法与人类设计方法相比的核心问题:它们的性能如何,以及它们的算法设计有何不同。为研究这些问题,本文引入了一项分析,该分析从人类设计的方法中推导特定任务的算法设计空间,将人类和智能体设计的方法映射到这些空间中,并在模块层面量化它们的算法差异。研究人员对广泛使用的LLM智能体在一组涵盖多种模态的代表性开放式AI任务上进行评估,从任务性能和与人类设计方法的算法差异两方面分析智能体设计的方法。实验结果显示,当前智能体偶尔能达到或超越人类的最新技术水平(SOTA)性能(72种配置中有10种),但这种成功无法在任务或智能体间可靠地推广。此外,96.8%的智能体设计方法属于人类推导的算法设计空间,主要是对人类设计方法中的算法选择进行重组,而近一半的方法与现有人类算法设计完全匹配。综合来看,这些发现表明,尽管当前智能体偶尔能达到或超越人类SOTA性能,但其算法设计仍处于人类推导的算法设计空间内,反映了算法选择的复用与重组。
英文摘要
Recent advances in LLM agents have made them increasingly capable of designing methods for complex AI tasks. This raises two central questions about agent-designed methods relative to human-designed methods: how well they perform, and how different their algorithmic designs are. To study these questions, this paper introduces an analysis that derives task-specific algorithmic design spaces from human-designed methods, maps both human- and agent-designed methods into these spaces, and quantifies their algorithmic differences at the module level. Widely used LLM agents are evaluated on a suite of representative, open-ended AI tasks spanning multiple modalities, and the methods they design are analyzed in terms of both task performance and algorithmic differences from human-designed methods. Experimental results show that current agents can occasionally match or surpass human state-of-the-art (SOTA) performance (10/72 configurations), but such success does not generalize reliably across tasks or agents. Moreover, 96.8% of agent-designed methods fall within human-derived algorithmic design spaces, largely recombining algorithmic choices found in human-designed methods, while nearly half exactly match an existing human algorithmic design. Taken together, these findings suggest that although current agents can occasionally match or surpass human SOTA performance, their algorithmic designs remain within human-derived algorithmic design spaces, reflecting the reuse and recombination of algorithmic choices.
发表机构
- University of Chinese Academy of Sciences(中国科学院大学)
- Northwestern University(西北大学)
机构由 AI 辅助整理,请以论文原文为准。