通过多模态Wasserstein重心绑定多种模态
Binding Multiple Modalities via Multimodal Wasserstein Barycenter
AI总结:
提出BaryBind方法,通过将特定模态传输至多模态Wasserstein重心并利用重心单纯形体积进行全局对齐,实现了多模态统一语义空间,在检索、分类等任务中表现优异且具鲁棒性和可扩展性。
AI中文摘要:
超越两种模态的多模态学习通常利用特定模态(如文本)来绑定其他模态。然而,如何建立一个更平衡的表示空间,既近似共享语义,又尊重n模态数据的整体几何结构,仍然具有挑战性。在这项工作中,我们提出了BaryBind,旨在将特定模态传输到在所有模态上优化的Wasserstein重心(WB),并引入体积对齐目标,以在WB嵌入周围建立统一的语义空间。具体来说,我们将特定模态投影到WB,这最小化了到多模态分布的平均Wasserstein距离,并作为后续对齐的锚点。然后,我们构建一个重心单纯形,其体积被用作以WB为中心的全局对齐的相似性度量。实验表明,BaryBind在文本-视频-音频检索、分类、视频问答和跨模态生成任务中取得了有竞争力的性能,并且在模态缺失情况下具有鲁棒性,且可扩展到三种以上模态。代码已在此https URL发布。
英文摘要:
Multimodal learning beyond two modalities commonly leverages a specific modality (e.g., text) to bind other modalities. However, how to establish a more balanced representation space that approximates shared semantics while respecting the holistic geometry of $n$-modal data remains challenging. In this work, we present BaryBind, which aims to transport the specific modality towards the Wasserstein barycenter (WB) optimized across all modalities and introduces a volumetric alignment objective to establish a unified semantic space around the WB embedding. Specifically, we project specific modalities to the WB, which minimizes the average Wasserstein distances to multimodal distributions and serves as the anchor for subsequent alignment. We then construct a barycenter simplex, whose volume is taken as a similarity metric for global alignment centered at the WB. Experiments show that BaryBind achieves competitive performance in text-video-audio retrieval, classification, videoQA, and cross-modal generation tasks, along with robustness under modality absence and scalability to more than three modalities. Code is released at https://github.com/xl-tang3/BaryBind.