arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

共享几何作为罗塞塔石碑:无需配对数据的跨模态对齐

Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data

Dominik Schnaus, Thomas Dagès, Daniel Cremers, Xi Wang, Phillip Isola

arXiv 2610.09411首次发表:更新:

发表机构

TU Munich; MCML; Ulm University; ETH Zurich; MIT(慕尼黑工业大学; 慕尼黑机器学习中心; 乌尔姆大学; 苏黎世联邦理工学院; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出Wasserstein Procrustes方法,利用共享几何结构在无需配对数据的情况下实现跨模态对齐,并证明其在少样本场景下优于现有方法,且能支持文本到图像生成。

AI 中文摘要

多模态表示能够实现零样本分类和检索,但对齐独立训练的模型通常需要大量配对数据。然而,柏拉图式表示假说表明,在不同模态上训练的模型可能自发地趋向于共享的表示几何。那么,我们是否真的需要配对示例来进行跨模态对齐?值得注意的是,我们证明对于粗略的跨模态对齐,配对示例并非必要。我们提出的简单Wasserstein Procrustes方法,通过粗略的几何初始化,在未见任何配对的情况下,通过估计单个正交映射来对齐两个不相交的嵌入集。跨数据集、模态和单模态模型,我们展示了一致地对齐独立训练的表示而无需配对是可行的,并且标准几何对齐指标能准确预测何时可行。尽管如此,我们自然能从配对示例中受益。在极少数配对的情况下,我们的方法显著优于现有方法,而在添加更多示例时与基于配对的方法保持竞争力。最后,我们展示了由此产生的对齐能够实现无需配对示例的文本到图像生成。这些结果表明,独立训练的模型通常共享足够的几何结构,以在很少或没有配对数据的情况下建立跨模态对应关系。

英文摘要

Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data. Yet, the Platonic Representation Hypothesis suggests that models trained on different modalities may converge spontaneously toward a shared representation geometry. But then, do we even need paired examples for cross-modal alignment? Remarkably, we show that paired examples are unnecessary for coarse cross-modal alignment. Our simple Wasserstein Procrustes method with a coarse geometric initialization aligns two disjoint embedding sets by estimating a single orthogonal map without seeing any pairs. Across datasets, modalities, and unimodal models, we show that we can consistently align independently trained representations without pairs, and standard geometric alignment metrics accurately predict when this is possible. Nevertheless, we can naturally benefit from paired examples. In the very few-pair regime, our method substantially outperforms existing ones, while staying competitive with pair-based methods with more added examples. Finally, we demonstrate that the resulting alignments can enable text-to-image generation without paired examples. These results show that independently trained models often share enough geometry to establish cross-modal correspondence with little or no paired data.

CommentsProject: https://dominik-schnaus.github.io/unpaired-rosetta/, Code: https://github.com/dominik-schnaus/unpaired-rosetta

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑