arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

支架分割掩盖了ADMET模型中的结构前沿失败

Beyond Scaffold Splits: Structural-Frontier Evaluation Reveals Hidden Failures in ADMET Models

Jiacheng Zheng, Chang Guo, Zixuan Wang, Xinyu Liu, Hao Chen

arXiv 2607.10729首次发表:更新:

发表机构

Ma Yinchu School of Economics, Tianjin University; Marine College, Shandong University(天津大学马寅初经济学院; 山东大学海洋学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究ADMET模型中支架分割掩盖结构前沿失败问题,引入无标签结构前沿分割并评估,对比多种分割方法,发现前沿分割会增大误差,多视图前沿风险外推法等测试结果不明确,强调分割构建和标签来源对评估的重要性。

AI 中文摘要

分子性质模型通常通过保留Bemis-Murcko支架来评估,但支架标识符只是化学陌生度的一种概念。我们引入了一种无标签的结构前沿分割,保留最稀疏和物理化学上最遥远的支架组,并在六个公共实验或策划的ADMET任务上进行评估。与具有相同无环分组的70/1...

英文摘要

Molecular property models are commonly evaluated by holding out Bemis-Murcko scaffolds, yet a scaffold identifier is only one notion of chemical unfamiliarity. We introduce a label-free structural-frontier split that reserves the sparsest and most physicochemically remote scaffold groups, and evaluate it on six public experimental or curated ADMET tasks. Against a 70/10/20 scaffold control with identical acyclic grouping, the frontier inflates equally weighted primary error with a taskwise median of 87.0% and a skew-sensitive mean of 130.3% (descriptive task/seed bootstrap interval, 52.1-246.0%). The mean falls to 75.9% once BBB is removed; that endpoint is the one whose score ranking inverts at the frontier. A message-passing graph-network control still shows a large gap (mean 82.8% over four tasks) and does not invert, so a low-capacity head does not explain the effect. We also test Multi-View Frontier Risk Extrapolation (MV-FREX), a count-adjusted tail-risk penalty over four molecular views, and treat it as a falsifiable probe. It changes normalized frontier error by only 0.16% relative to empirical risk minimization for the perceptron head (interval, -0.43-0.84%) and by -1.9% for the graph network; three fixed robust-penalty controls are likewise inconclusive. Against the published Lo-Hi and DataSAIL splitters, the frontier inflates error more on average, though no split is uniformly hardest. An audit of 31,561 marine natural products further shows that OOD status and agreement with legacy ADMET predictions depend on the molecular view, endpoint, and teacher coverage. Split construction and label provenance are important evaluation constraints in their own right, and the tested training penalties do not resolve the frontier failures we observe.

Comments15 pages, 4 figures, and 4 tables. Version 3 updates figures and tables caption in the main PDF

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑