发表机构
Department of Chemical Engineering, University of Michigan; Michigan Institute for Data & AI in Society, University of Michigan; Department of Statistics, University of Michigan(美国密歇根大学化学工程系; 美国密歇根大学数据与社会人工智能研究所; 美国密歇根大学统计系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何利用证据神经网络在测试时通过邻居融合优化分子性质预测,提出PG-EVIKAL方法,学习性质距离度量重排邻居,经实验验证该方法能降低RMSE、改善校准,还能在连续测定场景中优化预测,证明证据不确定性分解可用于测试时分子性质预测优化。
AI 中文摘要
一个经过训练的分子性质模型可以在测试时通过用最相似训练分子的测量标签校正每个预测来进行优化,这是一种无需重新训练的过程,我们称之为邻居融合;证据神经网络通过使用其偶然和认知不确定性对贝叶斯更新进行参数化,使其具有原则性。我们的主要贡献PG-EVIKAL,基于EVIKAL(标量卡尔曼滤波器)和GP-EVIKAL(处理相关邻居的高斯过程变体),学习一种性质距离度量,以便在融合前按性质相关性对结构相似的邻居重新排序。在16个分子数据集上评估,PG-EVIKAL相对于证据模型基线在其中14个数据集上降低了RMSE,中位数降低了19.4%,并改善了校准;在连续测定场景中,它进一步纳入新测量的分子,在不重新训练的情况下随着预测的到来进行优化。这项工作表明,证据不确定性分解不仅是一个校准目标,而且是一种可操作的推理资源,能够在测试时优化分子性质预测。
英文摘要
A trained molecular property model can be refined at test time by correcting each prediction with the measured labels of the most similar training molecules, a retraining-free procedure we call neighbor fusion; evidential neural networks make it principled by using their aleatoric and epistemic uncertainty to parameterize a Bayesian update. Our main contribution, PG-EVIKAL, learns a property-distance metric to re-rank structurally similar neighbors by their property relevance before fusion, building on EVIKAL (scalar Kalman filter) and GP-EVIKAL (Gaussian process variant handling correlated neighbors). Evaluated on 16 molecular datasets, PG-EVIKAL reduces RMSE relative to the evidential model baseline on 14 of them, with a median reduction of 19.4%, and improves calibration; in sequential-assay scenarios it further incorporates newly measured molecules, refining predictions as they arrive without retraining. This work demonstrates that evidential uncertainty decomposition is not merely a calibration objective but an actionable inference resource that enables test-time refinement of molecular property predictions.
Comments45 pages (18 main, 27 SI); 11 figures (7 main, 4 SI); 14 tables (0 main, 14 SI); 61 equations (15 main, 46 SI)