基于元学习在测定异质性下的少样本生物活性预测
Few-Shot Bioactivity Prediction with Meta-Learning under Assay Heterogeneity
浏览论文内容
中文总结 AI 辅助
针对测定异质性导致元学习性能下降的问题,提出MetaHeta框架,结合线性与精确注意力利用相关测定辅助数据,在ChEMBL和BindingDB上提升少样本生物活性预测与化合物排序。
中文摘要 AI 辅助
准确的生物活性预测是早期药物发现中的核心挑战,因为单个测定通常包含过少的测量数据,难以独立训练可靠的模型。元学习为这种少样本场景提供了一种原则性方法,但测定异质性可能限制其有效性。在此,我们验证了这一假设,并表明随着元训练任务变得更加异质,元学习性能会下降。为解决这一问题,我们引入了MetaHeta,一种元学习框架,通过利用相关测定的辅助数据来条件化预测,从而考虑测定异质性,其中相关性可根据现有测定信息灵活定义。MetaHeta的架构将大规模辅助数据集上的线性注意力与稀缺任务特定上下文上的精确注意力相结合,使得在扩展到前者时能够高效处理,同时不损害对后者的精确注意力。我们在ChEMBL和BindingDB的测定上展示了我们方法的优势,在回顾性贝叶斯优化中改善了少样本生物活性预测和下游化合物优先级排序。
英文摘要
Accurate bioactivity prediction is a central challenge in early-stage drug discovery, as individual assays often contain too few measurements to train reliable models independently. Meta-learning offers a principled approach to this few-shot setting, but assay heterogeneity may limit its effectiveness. Here, we test this hypothesis and show that meta-learning performance degrades as meta-training tasks become more heterogeneous. To address this, we introduce MetaHeta, a meta-learning framework that accounts for assay heterogeneity by conditioning predictions on auxiliary data from related assays, with relatedness defined flexibly from available assay information. The architecture of MetaHeta combines linear attention over large auxiliary datasets with exact attention over scarce task-specific context, enabling efficient scaling to the former without compromising exact attention over the latter. We demonstrate the benefits of our approach on assays from ChEMBL and BindingDB, improving few-shot bioactivity prediction and downstream compound prioritization in retrospective Bayesian optimization.
发表机构
- Helmholtz Munich(亥姆霍兹慕尼黑研究中心)
- Technical University of Munich(慕尼黑工业大学)
- Helmholtz AI(亥姆霍兹人工智能中心)
- Munich Center for Machine Learning(慕尼黑机器学习中心)
- University of Warsaw(华沙大学)
- University of Technology Nuremberg(纽伦堡工业大学)
- LMU Munich(慕尼黑大学)
机构由 AI 辅助整理,请以论文原文为准。