AI 中文总结
本文系统比较了不同信息假设下的最近邻检索变体,为MS/MS谱图指纹预测建立更严格的基线,以促进更严谨的基准测试和进展衡量。
AI 中文摘要
最近的研究表明,最近邻检索为从MS/MS谱图预测分子指纹提供了强大的基线,其中多种变体匹配或超越了当前的深度学习模型(Khoo和Barzilay,2026;Liu等人,2026;Gupta等人,2026)。重要的是,“最近邻”涵盖了一系列检索方法,这些方法在推理时可用的信息假设上有所不同。在本报告中,我们系统地比较了几种最近邻变体,并展示了这些不同的假设如何影响性能。我们的目标是建立更严格的基线,以实现更严谨的基准测试,并更好地衡量该领域的进展。
英文摘要
It has recently been shown that nearest-neighbour retrieval provides a strong baseline for molecular fingerprint prediction from MS/MS spectra, with several variants matching or outperforming current deep learning models (Khoo and Barzilay, 2026; Liu et al., 2026; Gupta et al., 2026). Importantly, "nearest neighbour" encompasses a family of retrieval methods that differ in the information assumed to be available at inference. In this report, we systematically compare several nearest-neighbour variants and show how these differing assumptions affect performance. Our goal is to establish stricter baselines that enable more rigorous benchmarking and better measure progress in this area.
Comments6 pages, 2 figures