发表机构
AI Laboratory for Molecular Engineering (AIME); Department of Computer Science and Engineering; Chalmers University of Technology & University of Gothenburg; Science for Life Laboratory (SciLifeLab)(分子工程人工智能实验室; 计算机科学与工程系; 查尔姆斯理工大学与哥德堡大学; 生命科学与技术实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对PROTAC渗透性预测数据稀缺问题,通过LLM提取流程将公开数据从31扩至87个,发现数据构成而非规模限制模型泛化,并指出报告实践改进方向。
AI 中文摘要
细胞渗透性是PROTAC开发的关键瓶颈,而可用于建模的公开数据稀缺且不一致。我们采用了一种专家参与循环的LLM提取工作流,从原始文献中挖掘PAMPA测量值,通过光学化学结构识别恢复仅含图像的化合物结构,并对每条记录进行人工验证,将公开记录从31个PROTAC扩展到87个。基于PROTAC-DB 3.0训练的岭回归模型在该资源内部达到R²=0.67,但在新提取的化学结构上性能崩溃(ρ=0.12),而基于新化合物训练的模型则能成功迁移回原有数据(ρ=0.80)。我们得出结论:当前公开记录的数据构成,而非数据集规模,限制了更通用模型的构建,并概述了在报告实践中需要改变的内容,以支持更好的数据驱动渗透性模型。
英文摘要
Cell permeability is a key bottleneck for PROTAC development, and public data available to model it is scarce and inconsistent. We adapt an expert-in-the-loop LLM extraction workflow to mine PAMPA measurements from the primary literature, recovering image-only structures by optical chemical structure recognition and hand-verifying every record, expanding the public record from 31 PROTACs to 87. Ridge models trained on PROTAC-DB 3.0 reach $R^2 = 0.67$ within that resource but collapse on the newly extracted chemistry ($ρ= 0.12$), while models trained on the new compounds transfer back successfully ($ρ= 0.80$). We conclude that the current composition of the published records, and not dataset size, is limiting the construction of more generalizable models, and we outline what would need to change in reporting practices for better data-driven permeability models.