可解释性设计的描述符组合在低数据分子测定中匹配2048维基础嵌入
Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays
- Harvey Mudd College(哈维穆德学院)
- Ersilia Open Source Initiative(Ersilia 开源计划)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出一种可解释性设计的描述符组合方法,在低数据分子测定中达到与2048维CheMeleon嵌入相当的预测性能,同时保持特征级可审计性。
AI中文摘要:
在低数据构效关系预测中,分子表示的选择可能比预测器的选择更重要,而表格基础模型加剧了这一效应。我们探究一组紧凑、语义命名的描述符块是否能在保持特征级可审计性的同时达到2048维CheMeleon嵌入的准确性,这意味着每个输入维度都带有模型名称和记录的训练来源。从固定的11维物理化学基础出发,我们仅利用标记上下文贪婪地拼接经过来源筛选的块。在九个ADME/Tox测定和50个评估单元中,基于限制在每种表示均覆盖的分子上的共同覆盖子集进行评分,该组合的平均测试AUC为0.762,而CheMeleon为0.764,Mordred为0.756。与CheMeleon的合并差距为+0.003 AUC(任务自举95%置信区间[-0.020, +0.030]),这满足了我们预先声明的合并平价门限,但不满足每测定门限。在25个上下文标签下,主要规则再次满足合并门限;在10个标签下则不满足。我们还报告了四个预先声明的候选选择规则,这些规则被证伪。冻结后对十个种子和三个先前未见测定的检查支持紧凑、可审计表示的合并竞争力;但等宽随机捆绑对照并未确立贪婪成员资格本身能增加准确性。测定水平差异仍未解决。
英文摘要:
In low-data structure-activity prediction, the choice of molecular representation can matter more than the choice of predictor, and tabular foundation models sharpen that effect. We ask whether a portfolio of compact, semantically named descriptor blocks can reach the accuracy of a 2048-dimensional CheMeleon embedding while staying auditable at the feature level, meaning that every input dimension carries a model name and a recorded training provenance. Starting from a fixed 11-dimensional physicochemical base, we greedily concatenate provenance-screened blocks using the labelled context alone. Across nine ADME/Tox assays and 50 evaluation cells, scored on common-coverage subsets restricted to the molecules that every representation covers, the portfolio reaches a mean test AUC of 0.762, against 0.764 for CheMeleon and 0.756 for Mordred. The pooled gap to CheMeleon is +0.003 AUC (task-bootstrap 95% CI [-0.020, +0.030]), which satisfies our predeclared pooled parity gate but not the per-assay gate. At 25 context labels the headline rule again satisfies the pooled gate; at 10 labels it does not. We also report four predeclared candidate-selection rules that we falsified. Post-freeze checks over ten seeds and three previously unseen assays support pooled competitiveness for compact, auditable representations; a same-width random-bundle control does not establish that greedy membership itself adds accuracy. Assay-level differences remain unresolved.