发表机构
Science for Life Laboratory, KTH Royal Institute of Technology; Biology Department, Brigham Young University; Munich Data Science Institute, Technical University of Munich(皇家理工学院生命科学实验室; 杨百翰大学生物系; 慕尼黑工业大学慕尼黑数据科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
dIon提出基于碎片化的不变性,利用前体与碎片离子关系进行自监督学习,提升肽段从头测序精度,超越全监督模型。
AI 中文摘要
我们提出了一种针对肽段串联质谱数据的新型不变性,从而解锁了自监督表示学习,提高了肽段的从头测序性能。这种不变性利用了前体性质(质量和电荷)与碎片离子证据之间的物理关系,无需肽段序列标签。我们引入了dIon,它改编了DINO框架,包含两个潜在预测任务,均用于恢复干净的教师表示:一个从谱图混合物中,使用前体作为选择查询;另一个从部分谱图中,隐藏前体信息。第一个任务将前体信息与碎片离子证据相关联;第二个任务防止表示坍缩到仅依赖该信息。机制探针支持这两种效应,消融实验表明完整目标函数表现最佳。在相同的端到端训练下,dIon初始化在保留的MassIVE-KB和Kingdoms测试集上,将从头测序肽段精度分别比从头训练提高了5.5和8.4个百分点,在使用更大的监督训练语料库时,分别提高了2.3和4.8个百分点。所得模型在多样化的多物种Kingdoms语料库上,在相同的贪婪解码协议下,超越了全监督的最先进从头测序模型。在没有肽段标签的情况下,与其他学习模型相比,dIon学习了强大的天然肽段相似性几何结构;在有限的肽段监督适应下,它在所有表示基准上实现了最佳的检索和配对判别性能。
英文摘要
We introduce a novel invariance for peptide tandem mass spectrometry data, unlocking self-supervised representation learning that improves de novo sequencing of peptides. This invariance exploits the physical relationship between precursor properties (mass and charge) and fragment-ion evidence, without requiring peptide sequence labels. We introduce dIon, which adapts the DINO framework with two latent prediction tasks, both recovering a clean teacher representation: one from a spectrum mixture, using the precursor as a selection query, and one from a partial spectrum with the precursor withheld. The first associates precursor information with fragment-ion evidence; the second prevents representational collapse onto that information alone. Mechanistic probes support both effects, and ablations show that the full objective performs best. Under identical end-to-end training, dIon initialization improves de novo peptide precision over training from scratch by 5.5 and 8.4 percentage points on the held-out MassIVE-KB and Kingdoms test sets, and by 2.3 and 4.8 percentage points with a larger supervised training corpus. The resulting models surpass fully supervised state-of-the-art de novo sequencing models on the diverse, multi-species Kingdoms corpus under the same greedy-decoding protocol. Without peptide labels, dIon learns strong native peptide-similarity geometry compared with other learned models; with limited peptide-supervised adaptation, it achieves the best retrieval and pair-discrimination performance across all representation benchmarks.
Comments37 pages, 10 figures, 29 tables. Code: https://github.com/statisticalbiotechnology/dIon