气味描述符语料库能测量与不能测量的内容:效价、衰减与公共记录的天花板
What an odour descriptor corpus can and cannot measure: valence, attenuation, and the ceiling of the public record
- Tesseract Academy(泰瑟拉克学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究审计四个气味描述符语料库,发现其不一致性无法通过汇总修复,且模型性能存在天花板;缺失的效价需直接测量,单次愉悦度评分显著提升预测,并发布交叉映射与定理。
AI中文摘要:
机器嗅觉在汇总的公共描述符语料库上进行训练,但共享的描述符词在不同语料库中是否测量相同的事物尚未被测试,也没有测试任何语料库所能测量的上限。我们审计了来自Pyrfume的四个语料库。以分子为条件,McNemar检验成为语料库效应的精确条件检验。语料库在描述符间异质地不一致($I^2 = 80\\%$),且随标签广度非均匀变化($z = 17.2$),因此没有单一的偏移量能修复汇总问题。中位四分相关一致性为0.795,而中位$\kappa$为0.413:来源大体上在哪些分子应获得某个词上一致,但在如何轻易应用该词上存在分歧。在109个具有可估计效应的描述符中,36个在ETS量表上显示出大的差异功能。相对于人类小组的可靠性,带有完整RDKit描述符块的Morgan指纹达到了可实现性的32.9\\%;从两个合并语料库中添加每个标签达到33.9\\%。这一差距不会因模型容量、编码选择、更多分子或更多词而缩小。缺失的方差是效价。每个分子一个愉悦度评分达到可实现性的54.6\\%(在独立的较旧仪器上为57.1\\%)。从描述符中恢复的效价($\rho = 0.457$)仅产生16.6\\%,因此必须直接测量。五名评分者超过了结构加上完整描述符记录;十五到二十名评分者达到饱和。我们发布了一个描述符交叉映射和二十个机器检查的定理。
英文摘要:
Machine olfaction trains on pooled public descriptor corpora, but whether a shared descriptor word measures the same thing across corpora has not been tested, nor has the ceiling of what any of them can measure. We audit four corpora from Pyrfume. Conditioning on the molecule makes McNemar's test the exact conditional test of the corpus effect. Corpora disagree heterogeneously across descriptors ($I^2 = 80\%$) and non-uniformly with labelling breadth ($z = 17.2$), so no single offset repairs pooling. Median tetrachoric agreement is 0.795 against median $κ$ of 0.413: sources largely concur on which molecules deserve a word and differ on how readily they apply it. Of 109 descriptors with an estimable effect, 36 show large differential functioning on the ETS scale. Against a human panel's reliability, Morgan fingerprints with the full RDKit descriptor block reach 32.9\% of achievable; adding every label from two merged corpora reaches 33.9\%. The gap does not close with model capacity, encoding choice, more molecules, or more words. The missing variance is valence. One pleasantness rating per molecule reaches 54.6\% of achievable (57.1\% on an independent older instrument). Valence recovered from descriptors ($ρ= 0.457$) yields only 16.6\%, so it must be measured. Five raters exceed structure plus the full descriptor record; fifteen to twenty saturate. We release a descriptor crosswalk and twenty machine-checked theorems.