arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

表型活性能否在没有实验读数的情况下被预测?

Can phenotypic activity be predicted without experimental readouts?

Télio Cropsal, Rocío Mercado

arXiv 2610.07997首次发表:更新:

发表机构

AI Laboratory for Molecular Engineering (AIME); Chalmers University of Technology; University of Gothenburg; Science for Life Laboratory (SciLifeLab)(AI分子工程实验室(AIME); 查尔姆斯理工大学; 哥德堡大学; 生命科学实验室(SciLifeLab))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估了预训练分子编码器(如CLOOME和CellCLIP)在无实验读数下预测表型活性的能力,通过控制泄漏和毒性混杂因素,发现其相比理化描述符无优势,强调需采用泄漏感知的混杂控制评估。

AI 中文摘要

分子编码器,如CLOOME和CellCLIP,通过对比学习在配对分子-形态学数据上进行预训练,已被提出作为表型预测的廉价替代方案,从而避免运行Cell Painting实验的必要性。我们在一种旨在控制两个可能夸大表观性能的混杂因素的协议下,对这些分子编码器的这一想法进行了评估:编码器自身预训练边界内的泄漏,以及表型活性与细胞毒性之间的相关性。我们在两个不同的Cell Painting筛选上测试了六种表示,包括一个与CLOOME输入和层数匹配的非预训练MLP对照,发现一旦这些混杂因素得到控制,预训练的分子编码器相对于普通理化描述符没有明显优势,并且毒性通常比表型活性更容易预测。我们的结果表明,在表型预训练编码器被信任为表型药物发现的替代方案之前,泄漏感知、混杂控制的评估应成为标准做法。

英文摘要

Molecular encoders contrastively pretrained on paired molecule-morphology data, such as CLOOME and CellCLIP, have been proposed as cheap surrogates for phenotypic prediction, avoiding the need to run a Cell Painting assay. We evaluate this idea for these molecular encoders under a protocol designed to control for two confounds that can inflate apparent performance: leakage across an encoder's own pretraining boundary, and the correlation between phenotypic activity and cytotoxicity. Testing six representations, including a non-pretrained MLP control matching CLOOME's input and layer count, on two distinct Cell Painting screens, we find that once these confounds are controlled for, the pretrained molecular encoders show no clear advantage over plain physicochemical descriptors, and that toxicity is generally easier to predict than phenotypic activity across representations. Our results suggest leakage-aware, confound-controlled evaluation should be standard practice before phenotype-pretrained encoders are trusted as surrogates for phenotypic drug discovery.

CommentsAccepted to the ML4Molecules: Agentic Systems for Molecular Sciences Workshop at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑