arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结构化潜在建模用于监督多模态信息分解

Structured Latent Modeling for Supervised Multimodal Information Decomposition

Wanting Huang, Sanvesh Srivastava, Weiran Wang

arXiv 2609.35502首次发表:更新:

发表机构

University of Iowa(爱荷华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种结构化潜在建模框架,通过中间层对比/掩蔽目标、逐源归一化流和低秩潜在变量模型,分解多模态信息为共享、模态特定和无关部分,提升预测性能。

AI 中文摘要

多模态预测依赖于多种形式的证据:跨模态重复的信息、特定于单一来源的线索,以及仅在输入被联合考虑时才出现的复杂跨模态依赖关系。虽然近期方法促进了更丰富的交互,但它们缺乏一种原则性的方式来在学习的连续表示中隔离这些与目标相关的贡献。我们引入了一个框架,在中间层应用对比或掩蔽目标,并结合逐源可逆归一化流和监督的低秩潜在变量模型。该架构明确地将联合分布分解为共享的任务相关变化、模态特定的预测变化和任务无关的依赖。借鉴先前的多模态学习假设,我们的方法评估模态如何独立地和联合地贡献于目标。最终,该框架将中间表示学习与结构化似然引导相结合,为刻画连续多模态交互提供了一个实用的潜在变量视角。在实证上,我们在多样化的多模态基准上展示了我们方法的有效性,显示出预测性能的稳健提升。

英文摘要

Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way to isolate these target-relative contributions within learned continuous representations. We introduce a framework that applies contrastive or masked objectives at intermediate layers, coupled with source-wise invertible normalizing flows and a supervised, low-rank latent variable model. This architecture explicitly factorizes the joint distribution into shared task-relevant variation, modality-specific predictive variation, and task-irrelevant dependence. Drawing connections to prior multimodal learning assumptions, our approach evaluates how modalities independently and jointly contribute to the target. Ultimately, this framework unites intermediate representation learning with structured likelihood-based guidance, offering a practical latent-variable lens for characterizing continuous multimodal interactions. Empirically, we demonstrate the effectiveness of our approach across diverse multimodal benchmarks, showing robust improvements in predictive performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑