发表机构
University of Chicago; Vector Institute; University of Toronto(芝加哥大学; 向量研究所; 多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出通过初始并行探索阶段初始化迭代优化,以提升大语言模型驱动的发现任务的成功率,并在多个框架和任务中验证了其一致收益。
AI 中文摘要
大语言模型(LLMs)已被用于通过使用提示LLM迭代优化目标的框架(harnesses)来新颖地发现算法、定理、药物及其他任务。在本工作中,我们研究了先前迭代的种群与最终发现成功之间的关系。我们推广了以往关于框架设计的工作,开发了一套名为“Modular”的12个框架,并表征了它们在5个不同发现任务上的性能,发现发现成功是脆弱的且对框架设计敏感。我们揭示了模式坍缩(mode collapse),其特征为迭代多样性的急剧下降,作为一个常见的失败模式。我们发现流行的最先进框架和旨在延长这种坍缩的多样性诱导框架干预措施产生了不一致的收益。我们的结果反而揭示早期发现的性能可预测最终成功。因此,我们提出了一种普遍适用的干预措施,该措施执行一个初始的并行探索阶段,以初始化后续的迭代优化。我们的方法在众多框架和目标应用中提供了一致的收益,确认了初始化在LLM驱动的发现中的重要性。
英文摘要
Large Language Models (LLMs) have been used for novel discovery of algorithms, theorems, drugs, and other tasks through the use of harnesses that prompt an LLM to iteratively optimize an objective. In this work, we study the relationship between the population of previous iterates and eventual discovery success. We generalize past work on harness design to develop a suite of 12 harnesses called 'Modular' and characterize their performance across 5 diverse discovery tasks, finding that discovery success is brittle and sensitive to harness design. We uncover mode collapse, characterized by a dramatic drop in the diversity of iterates, as a common failure mode. We find that popular state-of-the-art harnesses and diversity-inducing harness interventions, which aim to prolong this collapse, yield inconsistent gains. Our results instead uncover that the performance of early discoveries is predictive of eventual success. We therefore propose a universally applicable intervention that performs an initial stage of parallel exploration in order to initialize subsequent iterative optimization. Our method provides consistent gains across many harnesses and target applications, confirming the importance of initialization in LLM-driven discovery.