AI 中文总结
本研究提出基于集成共识的分子图半监督学习方法,可提升多场景下分子预测准确率,增强模型鲁棒性、降低校准误差,且单模型性能优于传统监督训练的完整集成。
AI 中文摘要
机器学习正通过加速性质预测、模拟及新分子与新材料发现,变革分子科学。该领域获取标注数据常成本高昂且耗时,而大量未标注分子数据易获取。标准半监督学习方法常依赖保留标签的数据增强,这在分子领域难以设计——微小改变可显著改变性质。本研究表明,依赖集成共识的半监督方法可在各类分子数据集、任务类型及图神经网络架构上提升预测准确率。研究发现,采用集成共识目标训练可增强模型鲁棒性,表现出类似知识蒸馏的效果;经此方式训练的集成中单个成员,在几乎所有情况下性能优于采用传统监督方式训练的完整集成。此外,此类半监督训练可降低校准误差。
英文摘要
Machine learning is transforming molecular sciences by accelerating property prediction, simulation, and the discovery of new molecules and materials. Acquiring labeled data in these domains is often costly and time-consuming, whereas large collections of unlabeled molecular data are readily available. Standard semi-supervised learning methods often rely on label-preserving augmentations, which are challenging to design in the molecular domain, where minor changes can drastically alter properties. In this work, we show that semi-supervised methods that rely on an ensemble consensus can boost predictive accuracy across a diverse range of molecular datasets, task types, and graph neural network architectures. We find that training with an ensemble consensus objective increases robustness in models and exhibits an effect similar to knowledge distillation; an individual member of an ensemble trained this way outperforms a full ensemble trained in a traditional supervised fashion in almost all cases. In addition, this type of semi-supervised training reduces calibration error.
CommentsICML