面向AI驱动的纳米药物发现:用于纳米自组装预测的基准测试与多模态学习框架
Towards AI-Driven Nanomedicine Discovery: A Benchmark and Multimodal Learning Framework for Nano Self-Assembly Prediction
浏览论文内容
中文总结 AI 辅助
针对纳米药物发现中湿实验室筛选成本高、现有方法缺乏标准化基准的问题,构建了NSA-Bench基准,开发了NSA-Net多模态框架,其在NSA-Bench上的ROC-AUC达0.9470±0.0112等,可支持方剂优化。
中文摘要 AI 辅助
纳米自组装可将分子组件组装成具有生物活性的纳米级结构。源自中药方剂的自组装纳米颗粒(NAPs)在抗肺癌治疗等应用中展现出纳米药物发现的巨大潜力。然而,当前的发现仍依赖成本高昂的湿实验室筛选,而现有机器学习方法缺乏标准化任务、有效的成对兼容性建模以及采用统一评估的公共基准。为解决这些局限,我们将纳米自组装(NSA)预测形式化为预测分子对之间自组装的二分类任务,进而构建了NSA-Bench——首个包含精选分子组合、实验条件、自组装标签及标准化评估协议的公共基准。我们还开发了NSA-Net,这是一种感知交互的多模态框架,它整合来自图拓扑、序列语义和物理化学描述符的互补分子证据,以学习用于自组装预测的分子对表示。在NSA-Bench上的大量实验表明,NSA-Net在小规模(Small)赛道上的ROC-AUC为0.9470±0.0112,在大规模(Large)赛道上为0.9492±0.0062;在小规模赛道上,它分别比最强的机器学习基线和图基线高出3.9和17.1个百分点。表示分析揭示了学习到的表示所捕获的与自组装预测相关的可解释分子特征。此外,一项NSA-Agent案例研究进一步证明,NSA-Net的预测如何通过感知实验条件的推理支持方剂优化。我们的代码可在此URL获取。
英文摘要
Nano self-assembly organizes molecular components into bioactive nanoscale structures. Self-assembled nanoparticles (NAPs) derived from Chinese herbal formulas and applications such as anti-lung-cancer therapy demonstrate the substantial potential of self-assembly for nanomedicine discovery. Yet discovery still relies on costly wet-lab screening, while existing machine learning approaches lack standardized tasks, effective pairwise compatibility modeling, and public benchmarks with unified evaluation. To address these limitations, we formalize NSA prediction as a binary classification task for predicting self-assembly between molecular pairs and then establish NSA-Bench, the first public benchmark with curated molecular combinations, experimental conditions, self-assembly labels, and standardized evaluation protocols. We further develop NSA-Net, an interaction-aware multimodal framework that integrates complementary molecular evidence from graph topology, sequence semantics, and physicochemical descriptors to learn molecular-pair representations for self-assembly prediction. Extensive experiments on NSA-Bench show that NSA-Net achieves a ROC-AUC of $0.9470\pm0.0112$ (Small) and $0.9492\pm0.0062$ (Large). On the Small track, it surpasses the strongest machine-learning and graph-based baselines by 3.9 and 17.1 percentage points, respectively. Representation analyses reveal interpretable molecular characteristics associated with self-assembly prediction captured by the learned representations. Moreover, an NSA-Agent case study further demonstrates how NSA-Net predictions can support formulation refinement through experimental-condition-aware reasoning. Our code is available at https://github.com/developer-hq/NSA-Net.
发表机构
- The School of Information Science and Technology, Beijing University of Technology(北京工业大学信息科学与工程学院)
- Artemisinin Research Center, China Academy of Chinese Medical Sciences(中国中医科学院青蒿素研究中心)
- College of Computer Science, Beijing University of Technology(北京工业大学计算机学院)
- School of Computer Science and Engineering, University of New South Wales(新南威尔士大学计算机科学与工程学)
- Institute of Automation, Chinese Academy of Sciences
机构由 AI 辅助整理,请以论文原文为准。