DoGMA:一种用于肿瘤学多组学对齐与多任务学习的中心法则引导基础模型
DoGMA: A Central-Dogma-Guided Foundation Model for Multi-Omics Alignment and Multi-Task Learning in Oncology
浏览论文内容
中文总结 AI 辅助
本研究提出中心法则引导的多组学基础模型DoGMA,通过Transformer-MoE架构结合定向注意力与掩码分层多组学重构预训练,在泛癌多组学下游任务中展现出优异性能,为多组学注意力机制设计提供新方向。
中文摘要 AI 辅助
注意力机制已在现代深度学习中得到广泛应用,许多现有多组学模型继承了其常规用法以实现无限制的双向交互。然而,生命的基本逻辑是具有方向性的,现有设计往往忽略了中心法则所暗示的方向性,这可能会限制跨异质性癌症、下游任务及不完整模态的迁移。在本研究中,我们提出了DoGMA,一种用于泛癌多组学分析的中心法则引导基础模型,认为稳健的迁移需要具有领域特定归纳偏置的表征。具体而言,我们在Transformer-MoE架构上构建该模型,其中定向注意力将多组学间的通信偏向中心法则信息流。我们进一步使用掩码分层多组学重构对模型进行预训练,以引导其学习与中心法则一致的交互。在包括癌症表征学习、生存预测和转移预测在内的各类下游任务中,DoGMA始终表现出强大的预测性能。 ablation实验及分析进一步表明,性能提升源于中心法则引导的定向注意力与基于重构的预训练之间的协同作用,二者共同促进了更符合生物学规律的多组学间信息交换。总体而言,DoGMA证明领域特定归纳偏置可提升多组学基础模型的稳健性与可迁移性,为多组学表征学习的注意力机制设计提供了新见解。
英文摘要
Attention mechanisms have been widely utilized in modern deep learning, and many existing multi-omics models inherit their conventional use to allow unrestricted bidirectional interactions. However, the fundamental logic of life is directional. Existing designs often overlook the directionality suggested by the central dogma, potentially limiting transfer across heterogeneous cancers, downstream tasks, and incomplete modality settings. In this work, we present DoGMA, a central-dogma-guided foundation model for pan-cancer multi-omics analysis, arguing that robust transfer requires representations with domain-specific inductive bias. Concretely, we build it on a Transformer-MoE architecture where directed attention biases inter-omics communication toward central-dogma information flow. We further pretrain our model with masked hierarchical omics reconstruction to guide it toward learning central-dogma-consistent interactions. Across diverse downstream tasks, including cancer representation learning, survival prediction, and metastasis prediction, DoGMA consistently demonstrates strong predictive performance. Ablations and analyses further suggest that the performance gains arise from the synergy between central-dogma-guided directed attention and reconstruction-based pretraining, which together promote more biologically consistent cross-omics information exchange. Overall, DoGMA demonstrates that domain-specific inductive biases can improve the robustness and transferability of multi-omics foundation models, offering new insights into the design of attention mechanisms for multi-omics representation learning.