DASyR-LLM:基于大语言模型的领域感知符号回归用于动力学模型发现
DASyR-LLM: Domain-Aware Symbolic Regression with LLMs for Kinetic Model Discovery
中文总结 AI 辅助
该研究提出了DASyR-LLM框架,将LLM嵌入迭代SR算法,通过领域知识减少41.7%-79.3%的动力学模型发现迭代次数,预测性能与现有方法相当。
中文摘要 AI 辅助
动力学模型发现是化学工程中的核心挑战,因为准确的速率表达式对于理解和控制化学及生物过程至关重要。符号回归(SR)已成为一种强大的数据驱动方法,用于识别可解释的动力学模型,但通常在没有领域知识的情况下运行,常探索出物理化学上不可行的模型。大语言模型(LLM)为将领域专业知识注入该搜索过程提供了有前景的途径。本文中,我们引入了一种LLM引导的SR框架,在迭代SR算法中嵌入LLM模块以实现自动化动力学模型发现。LLM在每次迭代中承担两个角色:(1)对最佳SR候选模型进行定性的物理化学评估;(2)基于SR生成的模型及嵌入的化学知识提出新的候选速率表达式。我们在四个复杂度递增的计算机模拟案例研究中对该框架进行评估,涵盖多相催化和生物过程系统。结果显示,与最先进的SR框架相比,LLM引导的框架将识别真实模型所需的迭代次数减少了41.7%至79.3%,且在超过一半的引导运行中,LLM直接提出了正确的模型结构。在实际场景中,每次迭代通常需要进行一次新的湿实验室实验,这意味着实验工作量大幅减少。在独立验证集上,两种方法的预测性能相当,所有案例研究的决定系数R²均大于0.98。消融研究表明,SR组件和LLM规模均对该性能有贡献,且缩小规模的LLM在很大程度上保留了发现效率。这些发现证明LLM可有效将领域知识注入科学模型发现过程,为实现完全自动化、领域感知的动力学建模 pipeline 铺平了道路。
英文摘要
Kinetic model discovery is a central challenge in chemical engineering, as accurate rate expressions are essential for understanding and controlling chemical and biological processes. Symbolic regression (SR) has emerged as a powerful data-driven approach for identifying interpretable kinetic models, but usually operates without domain knowledge, often exploring physicochemically implausible models. Large language models (LLMs) offer a promising avenue for injecting domain expertise into this search. Here, we introduce an LLM-guided SR framework, embedding an LLM module within an iterative SR algorithm for automated kinetic model discovery. The LLM performs two roles at each iteration: (1) a qualitative physicochemical critique of the best SR candidates, and (2) the proposal of new candidate rate expressions guided by the SR-generated models and embedded chemical knowledge. Our framework is evaluated on four in silico case studies of increasing complexity, spanning heterogeneous catalysis and bioprocess systems. Results show the LLM-guided framework reduces iterations to identify the ground-truth model by $41.7-79.3\%$ versus a state-of-the-art SR framework, with the LLM directly proposing the correct model structure in over half of the guided runs. In practical settings, where each iteration typically requires a new wet-lab experiment, this translates into a substantial reduction in experimental effort. Predictive performance on an independent validation set is equivalent between both approaches, with $R^2>0.98$ in all case studies. Ablation studies indicate that both the SR component and the LLM scale contribute to this performance, with a reduced-size LLM largely retaining discovery efficiency. These findings demonstrate that LLMs can effectively inject domain knowledge into scientific model discovery, paving the way toward fully automated, domain-aware kinetic modelling pipelines.