基于联合核熵 gromov-wasserstein 最优传输的多模态对齐
Multimodal Alignment Through Joint Kernel Entropic Gromov--Wasserstein Optimal Transport
- Northwestern University(西北大学)
- University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对跨模态配对数据稀缺的场景,提出 JK-EGW 框架实现多模态对齐,理论样本复杂度匹配标准最优传输,实验中在数据稀缺的预训练编码器嵌入对齐任务上性能优于基线。
AI中文摘要:
我们研究将多模态数据对齐到共享表示空间的问题,重点关注拥有强大预训练单模态编码器但跨模态配对数据稀缺的场景。我们提出一种保结构对齐框架:联合核熵 gromov-wasserstein 最优传输(JK-EGW),该框架通过最小化二次最优传输目标将多模态映射到公共潜在空间。JK-EGW 利用模态内部及跨模态的细粒度相似关系构建全局亲和核,而非依赖原始特征空间距离,还能显式控制潜在嵌入的几何结构与分布。理论层面,我们证明其参数样本复杂度率为 $n^{-1/2}$,与标准、熵及 gromov-wasserstein 最优传输的对应率匹配。算法层面,我们推导了一种可扩展的交替过程求解 JK-EGW,通过低秩核近似与变分提升结合熵最优传输(EOT)更新,该提升方案有效减轻了二次目标的负担,使我们能利用现有 EOT 求解器。实验上,我们聚焦数据稀缺场景下预训练编码器嵌入的事后对齐,结果显示,与现有对齐基线相比,所提方法实现了更优的多模态检索性能。
英文摘要:
We study the problem of aligning data from multiple modalities into a shared representation space, focusing on settings where strong pretrained unimodal encoders are available but cross-modal paired data are scarce. We propose a structure-preserving alignment framework, joint kernel entropic Gromov--Wasserstein Optimal Transport (JK-EGW), which maps multiple modalities into a common latent space by minimizing a quadratic optimal transport objective. JK-EGW leverages fine-grained similarity relationships within and across modalities to construct a global affinity kernel instead of relying on raw feature-space distances. Our framework naturally provides explicit control over the geometry and distribution of the latent embedding. On the theory side, we establish parametric sample complexity rate of $n^{-1/2}$, matching the corresponding rates for standard, entropic and Gromov--Wasserstein optimal transport. On the algorithmic side, we derive a scalable alternating procedure to solve JK-EGW with entropic optimal transport (EOT) updates through a low-rank kernel approximation and a variational lifting. This lifting scheme effectively relieves the burden of a quadratic objective, and allowing us to take the advantage of existing EOT solvers. Empirically, we focus on post-hoc alignment of embeddings from pretrained encoders in data-scarce regimes, and show that our proposed method achieves improved multimodal retrieval performance compared to existing alignment baselines.