arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GyroNovo:基于误差引导的片段填补与质量感知注意力用于从头肽段测序

GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for De Novo Peptide Sequencing

Abdellah El Mekki, Laks V. S. Lakshmanan, Muhammad Abdul-Mageed

arXiv 2609.30542首次发表:更新:

发表机构

The University of British Columbia(不列颠哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GyroNovo利用解码器误差引导填补并引入质量感知注意力,提升从头肽段测序精度,在NovoBench上显著超越现有方法。

AI 中文摘要

从串联质谱中进行从头肽段测序对于在不依赖参考数据库的情况下鉴定肽段至关重要。尽管深度学习取得了进展,但由于实验谱图通常稀疏、嘈杂且不完整,导致信息丰富的b-和y-离子片段未被观测到,因此准确测序仍然具有挑战性。现有方法尝试在自回归解码之前通过潜在空间填补来恢复这些缺失证据。然而,它们通常将填补视为固定的重建任务,而不考虑哪些缺失片段与解码器错误最相关。此外,现有的峰表示并未显式建模峰之间的质量差异,尽管这些差异具有根本重要性。我们引入了GyroNovo,一个具有两大贡献的框架。首先,我们利用训练期间观察到的解码器错误来调整填补目标,优先处理与频繁解码错误相关的片段。我们进一步利用解码器错误分布为每个谱图构建简单和困难的增强视图,使解码器能够在不同程度的谱图损坏和缺失片段严重性下学习。其次,我们通过使用旋转嵌入来编码谱峰之间的成对质量差异,将质量感知的归纳偏置引入自注意力机制。这些组件共同将缺失片段恢复与解码器行为对齐,同时显式纳入肽段碎裂背后的质量关系。在推理时,GyroNovo保持标准的编码器-填补器-解码器架构,既不需要额外输入,也不需要辅助搜索程序。在NovoBench上的实验显示,与最先进的基线相比,肽段水平精度提高了约9个百分点,氨基酸水平精度提高了7个百分点。代码:此https URL。

英文摘要

De novo peptide sequencing from tandem mass spectra is essential for identifying peptides without relying on reference databases. Despite advances in deep learning, accurate sequencing remains challenging because experimental spectra are often sparse, noisy, and incomplete, leaving informative b- and y-ion fragments unobserved. Existing methods attempt to recover this missing evidence via latent-space imputation before autoregressive decoding. However, they typically treat imputation as a fixed reconstruction task, without considering which missing fragments are most relevant to decoder errors. Moreover, existing peak representations do not explicitly model mass differences between peaks, despite their fundamental importance. We introduce GyroNovo, a framework with two main contributions. First, we use decoder errors observed during training to adapt the imputation objective, prioritizing fragments associated with frequent decoding errors. We further use the decoder error distribution to construct easy and hard augmented views of each spectrum, enabling the decoder to learn under varying degrees of spectral corruption and missing-fragment severity. Second, we introduce a mass-aware inductive bias into self-attention by using rotary embeddings to encode pairwise mass differences between spectral peaks. Together, these components align missing-fragment recovery with decoder behavior while explicitly incorporating the mass relationships that underlie peptide fragmentation. At inference time, GyroNovo retains a standard encoder-imputer-decoder architecture and requires neither additional inputs nor auxiliary search procedures. Experiments on NovoBench show gains of about 9 percentage points in peptide-level precision and 7 percentage points in amino-acid-level precision over the state-of-the-art baseline. Code: https://github.com/UBC-NLP/gyronovo.

CommentsCode available at https://github.com/UBC-NLP/gyronovo

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑