MultiLoReFT:通过低秩表示微调在多模态学习中解耦共享和特定模态子空间
MultiLoReFT: Decoupling Shared and Modality-Specific Subspaces in Multimodal Learning via Low-Rank Representation Fine-Tuning
浏览论文内容
中文总结 AI 辅助
研究针对多模态模型训练障碍,提出MultiLoReFT框架,通过低秩表示微调,将低秩适应扩展到多模态,学习可解释投影子空间解耦信息,能在多模态预测时揭示信息分布。
中文摘要 AI 辅助
现实世界中的感知和决策本质上是多模态的,整合跨模态的互补信号。然而,训练多模态模型面临两个主要障碍。一是收集大规模、对齐良好的配对多模态数据集通常不切实际;二是现有多模态表示常将跨模态共享信息与特定模态信息纠缠,阻碍解释性和控制。我们引入了MultiLoReFT,一种用于多模态学习的高效且可扩展的低秩表示微调框架。它将低秩适应扩展到多模态设置,学习可解释的投影子空间以解耦共享和特定模态信息。在模拟和现实基准测试中,它生成的表示支持多模态预测,同时明确揭示共享和特定模态信息如何跨模态分布。
英文摘要
Real-world perception and decision making are inherently multimodal, integrating complementary signals across modalities. However, training multimodal models faces two main obstacles. First, collecting large-scale, well-aligned paired multimodal datasets is often impractical, making end-to-end multimodal training difficult. Second, existing multimodal representations frequently entangle information shared across modalities with modality-specific information, hindering interpretability and control. We introduce MultiLoReFT, an efficient and scalable low-rank representation fine-tuning framework for multimodal learning with pretrained unimodal models. MultiLoReFT extends low-rank adaptation to the multimodal setting and learns interpretable projection subspaces that decouple shared and modality-specific information. Across simulated and real-world benchmarks, it produces representations that support multimodal prediction while explicitly revealing how shared and modality-specific information is distributed across modalities.