arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10857cs.LG

FiGuRO:面向多模态数据的本征维度估计

FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data

Viktoria Schuster, Sana Tonekaboni, Caroline Uhler

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出FiGuRO框架,用于在模型容量与超参数约束下估计单/多模态数据的本征维度,可实现共享与私有信息解耦,性能优于现有技术且可应用于单模态预训练模型。

中文摘要 AI 辅助

确定数据的复杂度即本征维度(ID),对于高效且可解释的表示学习至关重要,而在多模态场景下学习共享与私有信息的解耦表示时,这一任务尤其具有挑战性。现有技术存在关键缺陷:它们通常是静态的、单模态的,或者在对比方法的情况下,仅隐式地适配共享本征维度。我们提出Fidelity-Guided Rank Optimization(FiGuRO),这是一种在模型容量和超参数约束下近似单模态与多模态数据本征维度的框架。FiGuRO通过截断奇异值分解以及确定何时在哪个潜在空间中增大或减小维度的算法,学习低秩投影的维度。共享与私有信息的解耦是该优化的涌现属性,无需复杂的辅助损失函数。我们证明FiGuRO优于现有本征维度估计技术,且对超参数变化更具鲁棒性。在模拟数据和真实世界数据上,FiGuRO能捕获不同的本征维度尺度和变化的子空间比例,并成功分解共享与私有信息。此外,我们表明FiGuRO可应用于现代单模态预训练模型,实现高效的事后多模态表示解耦。

英文摘要

Determining the complexity, or Intrinsic Dimension (ID), of data is fundamental to efficient and interpretable representation learning. This is particularly challenging in multi-modal settings when trying to learn disentangled representations for shared and private information. Existing techniques leave a critical gap: they are often static, uni-modal, or in the case of contrastive methods, adapt only to the shared ID implicitly. We introduce Fidelity-Guided Rank Optimization (FiGuRO), a framework for approximating the ID of uni- and multi-modal data under constraints of model capacity and hyperparameters. FiGuRO learns the dimensions of low-rank projections using truncated singular value decomposition and an algorithm that determines when to reduce or increase dimension and in which latent space. Disentanglement of shared and private information arises as an emergent property of this optimization, eliminating the need for complex auxiliary loss functions. We demonstrate that FiGuRO outperforms existing ID estimation techniques and is more robust to hyperparameter changes. Across simulations and real-world data, FiGuRO captures distinct ID scales and varying subspace ratios, and decomposes shared and private information successfully. Furthermore, we show that FiGuRO can be applied to modern uni-modal pretrained models, enabling efficient, post-hoc disentanglement of multi-modal representations.

↑