DAIF:一种基于数据的近似消息传递多模态监督学习中间融合框架
DAIF: A Data-Driven Intermediate Fusion Framework for Multimodal Supervised Learning via Approximate Message Passing
浏览论文内容
中文总结 AI 辅助
本研究提出DAIF数据自适应中间融合框架,结合随机矩阵理论与非参数依赖度量,通过近似消息传递生成去噪特征,在模拟及两个多模态数据集上的预测任务中表现优于或媲美现有方法。
中文摘要 AI 辅助
多模态监督学习旨在利用多个异构数据源提升预测性能,核心挑战在于确定跨模态的融合粒度:过度融合可能放大噪声,而融合不足则无法利用跨模态依赖关系。现有方法依赖从早期融合到晚期融合的预定义融合架构,可能无法适配模态间的潜在依赖结构。我们提出DAIF,一种数据自适应中间融合框架,结合随机矩阵理论与非参数依赖度量直接从数据中学习融合结构。我们在贝叶斯多模态因子模型下开展研究,其中潜在因子的先验决定跨模态依赖关系。我们基于估计的跨模态依赖关系对模态进行聚类,随后执行聚类级别的经验贝叶斯先验估计。这些估计得到的先验被用于在近似消息传递(AMP)框架内构建去噪器,生成去噪后的低维特征,这些特征从相关模态中借用强度,同时保留模态特异性信号。生成的嵌入被用于下游监督预测。我们通过在不同依赖结构和信号 regime 下的模拟对该框架进行评估,与多种基准方法对比,并在两个多模态数据集上验证其实际效用,即三模态TEA-seq数据集(Swanson等人,2021)和TCGA-BRCA数据集(Goldman等人,2020)。在第一个案例中,我们预测T细胞分化标志物蛋白的表达水平;在第二个案例中,我们基于多模态信息分析患者生存预测。我们的方法在两个预测问题上与最先进技术相当或更优,证明其在不同监督学习任务中的通用性。
英文摘要
Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance. A central challenge is determining the fusion granularity across modalities: over-integration may amplify noise while under-integration fails to exploit cross-modal dependence. Existing approaches rely on pre-specified fusion architectures, from early to late fusion, that may not adapt to the underlying dependence structure among modalities. We propose DAIF, a data adaptive intermediate fusion framework that combines random matrix theory and non-parametric dependence measures to learn fusion structure directly from data. We operate under a Bayesian multimodal factor model where the prior on the latent factors determines the cross-modal dependence. Our method clusters modalities based on estimated intermodal dependence, then performs clusterwise empirical Bayes estimation of the priors. These estimated priors are used to construct denoisers within an approximate message passing (AMP) framework, yielding denoised low-dimensional features that borrow strength across related modalities while preserving modality-specific signal. The resulting embeddings are used for downstream supervised prediction. We evaluate the framework through simulations under varying dependence structures and signal regimes, comparing against several benchmark methods, and demonstrate its practical utility on two multimodal datasets, namely a trimodal TEA-seq dataset (Swanson et al., 2021) and TCGA-BRCA dataset (Goldman et al., 2020). In the first example, we predict the expression level of a T-cell differentiation marker protein and in the second case we analyze patient survival prediction based on multimodal information. Our method competes with or outperforms the state-of-the-art techniques in both prediction problems, demonstrating its versatility across diverse supervised learning tasks.
发表机构
- Ohio State University(俄亥俄州立大学)
- National University of Singapore(新加坡国立大学)
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。