arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越预测准确性:对人类嗅觉分子表征的可靠性感知审计

Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction

Kai Lun Huang, Wei Chieh Sun

arXiv 2607.24848首次发表:更新:

发表机构

California State University, Fullerton; University of Washington(加州州立大学富勒顿分校; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对人类嗅觉分子表征进行可靠性感知审计,涵盖多方面。通过多个数据集比较多种模型,发现人类评级几何可重复,模型与人类对齐弱,学习嵌入不一定优于传统表征,结果为分子编码器设边界,推动更广泛评估原则。

AI 中文摘要

预训练的分子编码器通常通过下游预测进行评估,但仅预测准确性并不能确定学习到的表征是否捕捉到可重复的科学结构、是否添加了超越强大传统基线的信息或是否能跨分布转移。我们对人类嗅觉的通用分子表征进行了可靠性感知审计,涵盖四个不同的方面:全局感知几何、超越化学的增量预测价值、跨数据集复制以及向未知成分的混合物转移。使用凯勒 - 沃斯霍尔和比尔林单分子评级数据集以及马二元混合物数据集,我们在身份控制和匹配评估下,将MoLFormer和ChemBERTa与RDKit描述符和摩根指纹进行比较。基于强度、愉悦度和熟悉度的人类三属性评级几何在参与者划分中是可重复的(中位数RSA为0.743和0.855),而模型与人类的对齐则弱得多(RSA为0.019 - 0.158)。在全局对齐中,学习到的嵌入并不总是优于传统表征,并且在任何一个单分子数据集中,MoLFormer都没有提供超越RDKit - 摩根组合基线的明确增量预测价值。人类几何在63个共享分子中显示出正但不完整的一致性(RSA为0.331;95%自举区间[0.204, 0.507])。在一种严格的未知成分混合物划分下,增量效应取决于结果和表征,所有区间都穿过零。这些结果为评估的通用分子编码器建立了经验边界,并推动了更广泛的评估原则:在科学领域中,应分别针对目标可靠性、结构对齐、增量信息、复制和跨分布转移来评估表征质量。

英文摘要

Pretrained molecular encoders are commonly evaluated through downstream prediction, but predictive accuracy alone does not establish that a learned representation captures reproducible scientific structure, adds information beyond strong conventional baselines, or transfers out of distribution. We present a reliability-aware audit of generic molecular representations for human olfaction across four distinct claims: global perceptual geometry, incremental predictive value beyond chemistry, cross-dataset replication, and mixture transfer to unseen components. Using the Keller-Vosshall and Bierling single-molecule rating datasets and the Ma binary-mixture dataset, we compare MoLFormer and ChemBERTa against RDKit descriptors and Morgan fingerprints under identity-controlled and matched evaluations. Human three-attribute rating geometry, based on intensity, pleasantness, and familiarity, is reproducible across participant splits (median RSA 0.743 and 0.855), whereas model-human alignment is substantially weaker (RSA 0.019-0.158). Learned embeddings do not consistently outperform conventional representations in global alignment, and MoLFormer provides no clear incremental predictive value beyond a combined RDKit-Morgan baseline in either single-molecule dataset. Human geometry shows positive but incomplete agreement across 63 shared molecules (RSA 0.331; 95% bootstrap interval [0.204, 0.507]). Under one strict unseen-component mixture split, incremental effects are outcome- and representation-dependent, with all intervals crossing zero. These results establish empirical boundaries for the evaluated generic molecular encoders and motivate a broader evaluation principle: representation quality in scientific domains should be assessed separately for target reliability, structural alignment, incremental information, replication, and out-of-distribution transfer.

Comments23 pages, 5 figures; includes supplementary material

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑