AI 中文总结
研究针对生物制造数据稀缺问题,将生物动力学ODE知识注入神经网络,比较数据级和架构级先验方法,二者在多数据集和微生物物种上均优于无先验基线且可替代,模拟预训练为生物过程深度学习提供简单高效方法。
AI 中文摘要
虽然深度学习加速了药物发现,但其对生物制造的影响较为有限,原因是数据稀缺。生物反应器实验成本高、耗时久且很少公开共享,各研究工作仅有少量实验数据。不过该领域有丰富先验知识,生物动力学常微分方程(ODE)模型已描述微生物生长数十年,但如何将此知识注入神经网络尚未系统研究。本文首次系统研究如何将ODE知识注入神经网络,比较了在模拟ODE曲线上预训练通用解码器的数据级先验与将ODE嵌入解码器的架构级先验。二者在11个数据集和7种微生物物种上均持续优于无先验基线,且二者可相互替代。模拟预训练为生物过程数据稀缺下的深度学习提供了简单且数据高效的方法。
英文摘要
While deep learning has accelerated drug discovery, its impact on biomanufacturing has been considerably more limited. The reason is data scarcity. Bioreactor experiments are high-cost, take days to weeks, and are rarely shared in public form, leaving each research work with only a handful of experiments. The domain itself, however, is rich in prior knowledge. Biokinetic ordinary differential equation (ODE) models have described microbial growth for decades, yet how to inject this knowledge into a neural network has not been studied systematically. We present the first systematic study of how to inject this ODE knowledge into a neural network, comparing a data-level prior that pre-trains a generic decoder on simulated ODE curves against an architecture-level prior that embeds the ODE inside the decoder. Both consistently outperform no-prior baselines across 11 datasets and 7 microbial species. Our central finding is that the two are substitutable. A generic decoder pre-trained on simulation matches a fully bio-structured decoder trained on real data. Simulation pre-training therefore offers a simple, data-efficient recipe for deep learning under bioprocess data scarcity.
CommentsAccepted at ICML 2026 AI for Science Workshop