多模态饱和的几何:黎曼VICReg
On the Geometry of Multimodal Saturation: Riemannian VICReg
浏览论文内容
中文总结 AI 辅助
该研究提出多模态饱和问题,即三模态自监督学习性能下降,并引入黎曼VICReg(R-VICReg)通过负曲率几何对齐视图,将第三模态有益概率从44.4%提升至64.4%。
中文摘要 AI 辅助
在自监督学习中,第三个模态应当提升或至少保持性能。在九个图像-文本-表格数据集上,我们发现它反而损害了性能:在VICReg下,三模态模型在55.6%的配对运行中表现不如其自身最佳的双模态子集。在SimSiam下,同样的失败发生在51.1%的配对运行中。我们将这种失败称为多模态饱和。我们提出,问题在于对齐几何。黎曼VICReg(R-VICReg)推广了经典VICReg:它通过可学习的负曲率乘积因子上的平方测地距离对齐视图,并在曲率消失时精确恢复VICReg。在相同的45次配对运行中,R-VICReg将第三模态有帮助的概率从44.4%提升到64.4%,增益集中在VICReg饱和的地方。
英文摘要
In self-supervised learning, a third modality should improve, or at least preserve, performance. Across nine image-text-tabular datasets, we show that it instead harms performance: the trimodal model underperforms its own best bimodal subset in 55.6% of paired runs under VICReg. The same failure occurs in 51.1% of paired runs under SimSiam. We call this failure multimodal saturation. We propose that the failure lies in the alignment geometry. Riemannian VICReg (R-VICReg) generalizes classical VICReg: it aligns views by squared geodesic distance on learnable negative-curvature product factors and recovers VICReg exactly as curvature vanishes. Over the same 45 paired runs, R-VICReg raises the probability that the third modality helps from 44.4% to 64.4%, with gains concentrated where VICReg saturates.
发表机构
- Protectline
- Télécom Paris(巴黎电信学院)
机构由 AI 辅助整理,请以论文原文为准。