arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作为研究者的AI智能体集群:进展、挑战与开放问题

AI Agent Swarms as Researchers: Progress, Challenges, and Open Questions

Sergey Gusev, David E. Bernal Neira

arXiv 2609.35719首次发表:更新:

发表机构

Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究让现成编码AI智能体集群仅靠范围说明和工具权限,在多领域快速产出大量研究成果,揭示智能体已能承担部分常规理论研究,同时提出了成果信任、学术评价等待解问题并发布实验数据。

AI 中文摘要

人工智能(AI)智能体——即连接工具并循环运行的语言模型——如今已能在极少监督下完成冗长的多步骤任务。我们给多组现成的编码智能体集群提供了从狭窄主题到整个领域的简短范围说明,赋予其文献访问权限与计算工具使用权,并下达一条固定指令:取得真实、正确、有用的进展,不要停止。我们未提供任何科学思路。数周内,这些智能体在优化理论与物理科学的五个领域产出了大量研究笔记、论文长度的草稿与形式化证明,并在第六个领域提出了未经测试的实验室实验方案。我们并未声称所有成果都正确或具有创新性,但它们绝非无用噪声:在我们目前已核查的内容中,未发现重大科学错误,且有若干结果已通过证明助手验证。智能体产出成果的速度快于我们的审阅速度;我们估计完整审阅这些内容需要数月时间。结合2026年两项广受讨论的借助智能体集群取得的数学研究成果,我们的实验表明,至少在我们测试的领域中,智能体已经能够承担大部分常规理论研究工作。这引发了我们目前无法解答的问题:当审阅而非成果产出成为稀缺资源时,如何信任研究结果;当人类投入仅为一条提示词时,学术荣誉与发表成果的意义何在;以及机构为何要聘用研究者而非直接购买计算时长;还有人们如何学习一个领域、在智能体成果的基础上推进研究,并在自己跟不上的研究中保持掌控力。研究机构尚未做好准备:模型的改进速度快于机构的变革速度,因此它们应当现在就决定如何应对智能体能力不断提升的局面。我们提出了初步观点,发布了截至2026年9月25日的智能体未编辑输出内容,并邀请读者在自己的研究领域重复该实验。

英文摘要

Artificial intelligence (AI) agents, language models connected to tools and run in a loop, can now carry out long, multi-step tasks with little supervision. We gave swarms of off-the-shelf coding agents a short statement of scope, from a narrow topic to a whole field, access to the literature and to computing tools, and one standing instruction: make real, correct, useful progress, and do not stop. We supplied no scientific ideas. Within weeks, the agents produced a large body of research notes, paper-length drafts, and formal proofs in five areas of optimization theory and physical science, and proposed untested laboratory experiments in a sixth. We do not claim that all of it is correct or new, but it is not noise: in what we have checked so far, we found no major scientific error, and several results are proved in a proof assistant. The agents produced results faster than we could review them; we estimate that a full review would take us months. Together with two widely discussed 2026 results in mathematics obtained with swarms, our runs suggest that agents can already do a large part of routine theoretical research, at least in areas that we experimented with. This raises questions we cannot yet answer: how to trust results when review, not production, is the scarce resource; what credit and publication counts mean when the human input is a prompt, and why institutions would pay researchers rather than buy computing time; and how people can learn a field, add to what agents do, and stay in control of research they cannot keep up with. Research institutions are not ready: models improve faster than institutions change, so they should decide now how to respond as capabilities increase. We offer tentative positions, release the agents' unedited output as of 25 September 2026, and invite readers to repeat the experiment in their own fields.

Comments78 pages, 1 figure, 3 tables. Research corpus: https://github.com/SECQUOIA/agent-swarm-research

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑