arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BarcodeMAE+:重新思考DNA条形码基础模型的掩码预训练与全局表示

BarcodeMAE+: Rethinking Masked Pretraining and Global Representations for DNA Barcode Foundation Models

Monireh Safari, Pablo Millan Arias, Scott C. Lowe, Lila Kari, Angel X. Chang, Graham W. Taylor

arXiv 2609.35877首次发表:更新:

发表机构

University of Waterloo; Vector Institute; Simon Fraser University; Alberta Machine Intelligence Institute (Amii); University of Guelph(滑铁卢大学; 向量研究所; 西蒙弗雷泽大学; 阿尔伯塔机器智能研究所(Amii); 圭尔夫大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BarcodeMAE+通过编码器-解码器掩码预训练和显式训练的全局[CLS]表示,提升了DNA条形码基础模型性能,最佳辅助目标因生物区域而异。

AI 中文摘要

许多DNA基础模型通过掩码序列的一部分并要求模型重建这些部分来进行预训练。标准的掩码预训练使编码器接触到推理时不存在的特殊[MASK]标记,从而在训练和下游使用之间产生不匹配。对于DNA条形码,显式全局序列表示(如[CLS]标记)的作用以及应如何训练它,也仍然理解不足。我们引入了BarcodeMAE+,并研究了模型架构、全局[CLS]表示以及跨节肢动物COI(BIOSCAN-5M)和真菌ITS(UNITE+INSD)条形码的辅助预训练目标。在这两个条形码区域中,编码器-解码器MAE-LM架构在几乎所有评估配置中都优于其匹配的仅编码器对应架构,支持MAE-LM作为DNA条形码基础模型的有效架构设计。经过训练的全局[CLS]表示提供了显著的额外增益:在BIOSCAN-5M上,[CLS]准确率从无辅助目标时的47.53%提高到使用交叉熵属分类时的80.65%。最佳辅助目标因区域而异:交叉熵在BIOSCAN-5M上表现最佳,而配对同属分类在UNITE+INSD上表现最佳,在酵母菌上达到73.19%,在丝状真菌上达到63.07%。BarcodeMAE+在BIOSCAN-5M上优于已发表的DNA基础模型基线,并在评估的UNITE+INSD基线中使用冻结编码器表示实现了最高的酵母菌准确率。相似性加权softmax KNN投票进一步稳定了准确率,随着邻域大小的增加而提高。总体而言,编码器-解码器掩码预训练和显式训练的全局表示是DNA条形码基础模型的强有力设计选择,而学习该表示的最佳目标取决于生物学领域。

英文摘要

Many DNA foundation models are pretrained by masking parts of a sequence and asking the model to reconstruct them. Standard masked pretraining exposes the encoder to special [MASK] tokens that are absent at inference, creating a mismatch between training and downstream use. The role of an explicit global sequence representation such as a [CLS] token and how it should be trained also remain poorly understood for DNA barcodes. We introduce BarcodeMAE+ and study model architecture, global [CLS] representation, and auxiliary pretraining objectives across arthropod COI (BIOSCAN-5M) and fungal ITS (UNITE+INSD) barcodes. Across both barcode regions, the encoder-decoder MAE-LM architecture outperforms its matched encoder-only counterpart in nearly all evaluated configurations, supporting MAE-LM as an effective architectural design for DNA barcode foundation models. A trained global [CLS] representation provides substantial additional gains: on BIOSCAN-5M, [CLS] accuracy increases from 47.53% without an auxiliary objective to 80.65% with cross-entropy genus classification. The best auxiliary objective is region-dependent: cross-entropy performs best on BIOSCAN-5M, whereas pairwise same-genus classification performs best on UNITE+INSD, reaching 73.19% on Yeast and 63.07% on Filamentous Fungi. BarcodeMAE+ outperforms published DNA foundation model baselines on BIOSCAN-5M and achieves the highest Yeast accuracy among the evaluated UNITE+INSD baselines using frozen encoder representations. Similarity-weighted softmax KNN voting further stabilizes accuracy as neighbourhood size increases. Overall, encoder-decoder masked pretraining and an explicitly trained global representation are strong design choices for DNA barcode foundation models, while the optimal objective for learning that representation depends on the biological domain.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑