发表机构
University of Cambridge; Max Planck Institute for Intelligent Systems(剑桥大学; 马克斯·普朗克智能系统研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ARO提出一种对齐表示学习方法,从多组学癌症数据中学习表示以重建缺失和未配对模态,在有限数据下实现高重建精度(MSE 0.15),并支持下游分类任务,提供经济高效的多组学分析方案。
AI 中文摘要
功能分子检测的高成本、计算生物学中缺失模态和不匹配样本的普遍存在,为全面的多组学分析(这对于捕获和推理分子、细胞、组织和生物体至关重要)设置了重大障碍。本工作提出了一种模型,该模型从多组学癌症数据中学习有意义的表示,以支持缺失和未配对模态的重建。与日益复杂的大型模型(例如基础模型)相反,ARO优先考虑在有限或不完整数据设置中的实际适用性。ARO能最优地重建缺失模态(在未掩蔽设置下,验证和测试数据的均方误差为0.15),其学习到的潜在嵌入支持下游癌症分类任务。我们的研究结果表明,将不同的分子层作为一个集成系统进行分析,提供了一种可靠且经济高效的方法,减少了对大规模实验测试的依赖,同时仍支持有限数据设置下的多组学探索。
英文摘要
The high cost of functional molecular assays, and prevalence of missing modalities and unmatched samples in computational biology, create significant barriers to comprehensive multi-omic profiling, essential for capturing and reasoning over molecules, cells, tissues, and organisms. This work proposes a model that learns meaningful representations from multi-omics cancer data supporting the reconstruction of missing and unpaired modalities. Contrary to increasingly complex, larger models, e.g. Foundation Models (FMs), ARO prioritizes practical applicability in limited or incomplete data settings. ARO optimally reconstructs missing modalities (MSE of $0.15$ on the validation and test data in the Unmasked settings), with its learned latent embeddings enabling a downstream cancer classification task. Our findings indicate that analyzing diverse molecular layers as a single integrated system offers a reliable and cost-efficient approach, reducing dependence on large-scale experimental testing, while still supporting multi-omic exploration in limited data settings.
CommentsProceedings of the ICML 2026 3rd Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences, Seoul, Korea