arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

血脑屏障通透性的特征空间选择与异质性效应估计:从随机森林到广义随机森林的流程

Feature Space Selection and Heterogeneous Effect Estimation for Blood-Brain Barrier Permeability: A Random Forest to the Generalized Random Forest Pipeline

Tshemollo Rapolai, Seite Makgai, Mohammad Arashi

arXiv 2609.29076首次发表:更新:

发表机构

University of Pretoria; Ferdowsi University of Mashhad(比勒陀利亚大学; 马什哈德菲尔多西大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究利用MoleculeNet BBBP数据集,通过系统消融特征空间并应用广义随机森林与双重机器学习,发现特征表示与算法耦合影响预测性能,且校正混杂后LogP-BBB关联的异质性证据不足。

AI 中文摘要

预测血脑屏障(BBB)通透性对于中枢神经系统药物发现至关重要。本研究使用MoleculeNet BBBP数据集(n=2039),系统性地消融分子特征空间,以将特征化与模型架构分离开来。我们评估了三种特征族(Morgan指纹、RDKit理化描述符、SMILES二元组)在四种学习算法上的表现。结果表明,预测性能同时依赖于特征表示和算法。使用组合特征的动态随机森林达到了最高的平均AUC(0.970,95%置信区间:0.963-0.977)。其次,这种最优表示使得能够利用广义随机森林探索性地估计分子结构与BBB通透性之间的异质性关联。通过从LogP中位数分割构建伪处理,我们应用双重/去偏机器学习来考虑混杂因素。正交化显著减弱了朴素因果森林检测到的异质性;在错误发现率校正后,没有条件效应仍然显著(最小调整后p=0.082)。此外,正交化的特征重要性转向了SMILES二元组中的残余结构信息。最终,一旦正确考虑了观察到的混杂因素,关于LogP-BBB关联在化学空间中系统性变化的证据不足。这强调了特征表示和模型架构是耦合的设计选择,并且未正交化的因果森林有高估真实处理效应异质性的风险。

英文摘要

Predicting blood-brain barrier (BBB) permeability is critical for central nervous system drug discovery. Using the MoleculeNet BBBP dataset (n = 2039), this study systematically ablates molecular feature spaces to isolate featurisation from model architecture. We evaluate three feature families (Morgan fingerprints, RDKit physicochemical descriptors, SMILES bigrams) across four learning algorithms. Results demonstrate that predictive performance depends jointly on feature representation and algorithm. Dynamic Random Forest using combined features achieved the highest mean AUC (0.970, 95% CI: 0.963-0.977). Second, this optimal representation enables exploratory estimation of heterogeneous associations between molecular structure and BBB permeability using Generalized Random Forests. Constructing a pseudo-treatment from a LogP median split, we applied double/debiased machine learning to account for confounding. Orthogonalization substantially attenuates the heterogeneity detected by naive causal forests; no conditional effects remained significant after false discovery rate correction (smallest adjusted p = 0.082). Furthermore, orthogonalized feature importance shifted toward residual structural information in SMILES bigrams. Ultimately, once observed confounding is properly accounted for, evidence that LogP-BBB associations vary systematically across chemical space is insufficient. This underscores that feature representation and model architecture are coupled design choices, and that unorthogonalized causal forests risk overstating genuine treatment effect heterogeneity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑