arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用生物动力学知识先验进行数据稀缺的生物过程建模

Leveraging Biokinetic Knowledge Priors for Data-Scarce Bioprocess Modeling

Kyunghoon Hur, Eunjung Jeon, Hyun Woo Kim, Gyubok Lee, Seongjun Yang

arXiv 2607.20539首次发表:更新:

AI 中文总结

研究针对生物制造数据稀缺问题,将生物动力学ODE知识注入神经网络,比较数据级和架构级先验方法,二者在多数据集和微生物物种上均优于无先验基线且可替代,模拟预训练为生物过程深度学习提供简单高效方法。

AI 中文摘要

虽然深度学习加速了药物发现,但其对生物制造的影响较为有限,原因是数据稀缺。生物反应器实验成本高、耗时久且很少公开共享,各研究工作仅有少量实验数据。不过该领域有丰富先验知识,生物动力学常微分方程(ODE)模型已描述微生物生长数十年,但如何将此知识注入神经网络尚未系统研究。本文首次系统研究如何将ODE知识注入神经网络,比较了在模拟ODE曲线上预训练通用解码器的数据级先验与将ODE嵌入解码器的架构级先验。二者在11个数据集和7种微生物物种上均持续优于无先验基线,且二者可相互替代。模拟预训练为生物过程数据稀缺下的深度学习提供了简单且数据高效的方法。

英文摘要

While deep learning has accelerated drug discovery, its impact on biomanufacturing has been considerably more limited. The reason is data scarcity. Bioreactor experiments are high-cost, take days to weeks, and are rarely shared in public form, leaving each research work with only a handful of experiments. The domain itself, however, is rich in prior knowledge. Biokinetic ordinary differential equation (ODE) models have described microbial growth for decades, yet how to inject this knowledge into a neural network has not been studied systematically. We present the first systematic study of how to inject this ODE knowledge into a neural network, comparing a data-level prior that pre-trains a generic decoder on simulated ODE curves against an architecture-level prior that embeds the ODE inside the decoder. Both consistently outperform no-prior baselines across 11 datasets and 7 microbial species. Our central finding is that the two are substitutable. A generic decoder pre-trained on simulation matches a fully bio-structured decoder trained on real data. Simulation pre-training therefore offers a simple, data-efficient recipe for deep learning under bioprocess data scarcity.

CommentsAccepted at ICML 2026 AI for Science Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑