发表机构
Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文回顾了ML在定向进化领域进展滞后的原因,指出MLDE与定向进化的目标脱节及忽略DNA合成成本是关键,还提及例外工作并认为二者可兼顾。
AI 中文摘要
过去五年多,机器学习(ML)的进展已改变了许多蛋白质工程学科,但定向进化领域并非如此。本文回顾了之前合著的一篇观点文章,探讨了我认为出现这种情况的原因,指出机器学习辅助定向进化(MLDE)研究者的目标是“识别最优蛋白质”,而更广泛的定向进化的目标是“在时间和资源约束下识别足够的蛋白质”,这种目标脱节是主要原因。例如,我强调几乎所有当前的MLDE方法都忽略了DNA合成的成本,导致无论底层模型能力如何,策略的实际适用性都有限。最后,我讨论了这一总体趋势中的例外近期工作,并强调过去五年在机器学习辅助蛋白质工程方面的努力与提出的MLDE目标重新定位并非相互排斥。
英文摘要
The last five-plus years have seen many protein engineering disciplines transformed by advances in machine learning (ML), but the same cannot be said for directed evolution. Reflecting on a previously co-authored perspective, I discuss why I believe this to be the case, arguing that a disconnect between the goals of machine-learning-assisted directed evolution (MLDE) researchers--"identify an optimal protein"--and the goals of directed evolution more broadly--"identify a sufficient protein given time and resource constraints"--is a principal culprit. As an example, I highlight how nearly all current MLDE methods neglect to account for the cost of DNA synthesis, resulting in strategies that have limited practical applicability regardless of the underlying models' capabilities. I close by discussing recent works that are exceptions to this overarching trend, and emphasize that the last five years of efforts in ML-assisted protein engineering and the prescribed reframe of MLDE objectives need not be mutually exclusive.