AI 中文总结
该研究提出The Little Scientist框架,让LLM智能体遵循科学方法迭代探索,在蛋白质适应性预测、DNA基序发现问题上发现了优于现有方案的新算法与集成策略。
AI 中文摘要
当你向基于大语言模型(LLM)的智能体传授科学方法时,会发生什么?动机:科学发现源于假设、实现、实证测试与反馈的循环过程,这一过程能否实现自动化?我们从科学方法的视角入手,通过有序迭代的方式让基于LLM的智能体完成该过程的每一步,以此实现自动化算法设计。结果:我们提出了The Little Scientist框架,其中“科学家智能体”在评估环境中运行,该环境会对其代码进行基准测试并返回结构化的单实例诊断结果。当科学家智能体陷入局部最优时,“库恩智能体”会注入一个范式转变猜想,并搭配跨学科灵感,促使智能体探索LLM潜在空间的不同区域。我们在两个需要截然不同发现模式的问题上验证了该框架:对于蛋白质适应性预测,科学家智能体发现了Delta V,这是一种集成校准策略,在ProteinGym DMS替换零样本排行榜的全部5项官方评估指标中排名第一,在217项DMS检测中,其平均斯皮尔曼相关系数比第二名模型VenusREM高出+0.033;对于DNA基序发现,科学家智能体从零编写了算法DALE(Dual-seed Algorithm for Latent Enumeration),在132个ENCODE转录因子上,其性能优于STREME(MEME Suite中的默认工具),平均AUROC为0.842,而STREME为0.803,威尔科克森检验p值小于10^{-6},且运行速度快11倍。这表明该框架能够生成真正新颖的算法,而非仅优化现有组件。综合这些结果可知,遵循科学方法的LLM智能体可发现优于现有解决方案的新算法与新集成策略。整个研究项目在无GPU的单台虚拟机上消耗了7.04亿个token。
英文摘要
What happens when you teach an LLM-based agent the scientific method? Motivation: Scientific discovery emerges from cycles of hypothesis, implementation, empirical testing, and feedback. Can this process be automated? We approach automated algorithm design through the lens of the scientific method, where an LLM-based agent goes through each step of the process in an ordered, iterative fashion. Results: We present The Little Scientist, a framework in which a "Scientist agent" works inside an evaluation environment that benchmarks its code and returns structured per-instance diagnostics. When the Scientist plateaus at a local optimum, a "Kuhn agent" injects a paradigm-shifting conjecture paired with a cross-disciplinary inspiration, forcing exploration of a different region of the LLM's latent space. We demonstrate the framework on two problems that require fundamentally different modes of discovery. For protein fitness prediction, the Scientist discovered Delta V, an ensemble calibration strategy that ranks first on the ProteinGym DMS Substitutions Zero-Shot leaderboard across all five official evaluation metrics, exceeding the #2 model (VenusREM) by +0.033 mean Spearman correlation across 217 DMS assays. For DNA motif discovery, the Scientist wrote an algorithm from scratch--DALE (Dual-seed Algorithm for Latent Enumeration)--that outperforms STREME (the default in the MEME Suite) across 132 ENCODE transcription factors (mean AUROC 0.842 vs. 0.803, Wilcoxon p < 10^{-6}) while running 11x faster. This demonstrates that the framework can produce genuinely novel algorithms, not just optimize existing components. Together, these results show that an LLM agent stepping through the scientific method can discover both new algorithms and new ensemble strategies that outperform prior solutions. The entire research program consumed 704M tokens on a single virtual machine with no GPUs
CommentsCode: https://github.com/travis42/little-scientist-dale and https://github.com/travis42/little-scientist-delta-v (Apache 2.0)