arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06096cs.LGcs.AI

多模态饱和的几何:黎曼VICReg

On the Geometry of Multimodal Saturation: Riemannian VICReg

Nessim Ben Abbes, Duc Han Le, Sabri Mtibaa, Van-Tam Nguyen

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出多模态饱和问题,即三模态自监督学习性能下降,并引入黎曼VICReg(R-VICReg)通过负曲率几何对齐视图,将第三模态有益概率从44.4%提升至64.4%。

中文摘要 AI 辅助

在自监督学习中,第三个模态应当提升或至少保持性能。在九个图像-文本-表格数据集上,我们发现它反而损害了性能:在VICReg下,三模态模型在55.6%的配对运行中表现不如其自身最佳的双模态子集。在SimSiam下,同样的失败发生在51.1%的配对运行中。我们将这种失败称为多模态饱和。我们提出,问题在于对齐几何。黎曼VICReg(R-VICReg)推广了经典VICReg:它通过可学习的负曲率乘积因子上的平方测地距离对齐视图,并在曲率消失时精确恢复VICReg。在相同的45次配对运行中,R-VICReg将第三模态有帮助的概率从44.4%提升到64.4%,增益集中在VICReg饱和的地方。

英文摘要

In self-supervised learning, a third modality should improve, or at least preserve, performance. Across nine image-text-tabular datasets, we show that it instead harms performance: the trimodal model underperforms its own best bimodal subset in 55.6% of paired runs under VICReg. The same failure occurs in 51.1% of paired runs under SimSiam. We call this failure multimodal saturation. We propose that the failure lies in the alignment geometry. Riemannian VICReg (R-VICReg) generalizes classical VICReg: it aligns views by squared geodesic distance on learnable negative-curvature product factors and recovers VICReg exactly as curvature vanishes. Over the same 45 paired runs, R-VICReg raises the probability that the third modality helps from 44.4% to 64.4%, with gains concentrated where VICReg saturates.

发表机构

  • Protectline
  • Télécom Paris(巴黎电信学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑