生物化学领域中不确定性下的LLM序贯决策
LLM sequential decision making under uncertainty in biochemical domains
中文总结 AI 辅助
本研究在贝叶斯优化框架下对五个前沿LLM在七个生化数据集上的决策行为进行基准测试,发现其决策策略不可见且存在上下文粘性导致的能力差距,需解耦先验与数据以实现有效探索。
中文摘要 AI 辅助
大型语言模型(LLMs)正越来越多地被用于推动科学发现。在信任它们于紧张的实验预算下设计实验之前,理解LLMs如何从新数据和文献记忆中做出决策至关重要。然而,它们的决策策略在当前用于评估研究智能体的性能分数中是不可见的。在此,我们在贝叶斯优化框架下,对五个前沿LLM与已发表的统计基线在七个组合数据集上进行了基准测试,这些数据集涵盖蛋白质工程、反应优化、分子设计、肽自组装和催化领域。性能与模型信念和行为的直接测量相结合,实现了高分辨率的行为分析。一种逐步剥离上下文的提示消融实验将记忆与化学推理以及裸类别优化区分开来。先前的化学知识在期望上有所帮助,但方差很大,有时甚至损害性能。没有任何配置在测试中能决定性地超越跨领域的平均统计基线。信念-移动和Martingale诊断(此处针对一种将理性智能体误标为不理性的测量噪声偏差进行了校正)表明,在本研究涉及的上下文中,模型对传入数据反应过度,而非固守其先验。有趣的是,尽管LLM的行为是剥削性的,但模型真诚地意图探索并始终如一地按该意图行事。这种失败是由于上下文粘性导致的能力差距。移除上下文历史可恢复探索行为,表明为了有效实现LLM驱动的发现,必须将先验与数据解耦。
英文摘要
Large language models (LLMs) are increasingly used to drive scientific discovery. Understanding how LLMs make decisions from new data and memory of the literature is vital before trusting them to design experiments under tight experimental budgets. However, their decision strategies are invisible in the current performance scores used to evaluate research agents. Here, we benchmark five frontier LLMs in a Bayesian Optimization setting against published statistical baselines on seven combinatorial datasets spanning protein engineering, reaction optimization, molecular design, peptide self-assembly, and catalysis. Performance is paired with direct measurements of model beliefs and actions, enabling highly resolved behavior analysis. A prompt ablation that progressively strips context separates memorization from chemical reasoning and from bare categorical optimization. Prior chemical knowledge helps in expectation, but with high variance and occasionally even harms performance. No configuration tested decisively beats a mean statistical baseline across domains. Belief-movement and Martingale diagnostics, corrected here for a measurement-noise bias that mislabels rational agents as irrational, show that models overreact to incoming data rather than entrenching on their priors in the contexts studied here. Interestingly, while LLM actions are exploitative, models sincerely intend to explore and consistently act on that intent. This failure is a competence gap arising from context-stickiness. Removing in-context history restores exploration, indicating that priors and data must be decoupled to achieve effective LLM-driven discovery.