CodeBind: 一种用于多模态对齐的解耦表示学习框架
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
- Visual AI Lab, The University of Hong Kong(视觉人工智能实验室,香港大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
CodeBind通过统一的组合代码本设计优化多模态表示空间,解决了传统方法在跨模态信息差异和数据稀缺导致的对齐空间不足问题,实现了多模态分类和检索任务中的最佳性能。
AI中文摘要:
多模态表示对齐对于大语言模型和机器人至关重要。传统方法常受到跨模态信息差异和数据稀缺的限制,导致对齐空间不优,忽略了模态特有的特征。我们提出了CodeBind,一种通过模态共享-特定代码本设计优化多模态表示空间的框架。通过逐步对齐目标和连接模态,CodeBind避免了需要完全配对数据的需要。不同于传统硬对齐,CodeBind将特征分解为共享组件以实现语义一致性,以及特定组件以捕捉模态特有的细节。这种设计利用了组合向量量化方案,其中共享代码本弥合模态差距,而模态特定代码本通过防止主导模态压制其他模态来缓解表示偏差。在九种模态(文本、图像、视频、音频、深度、热成像、触觉、3D点云、EEG)上验证,CodeBind在多模态分类和检索任务中实现了最先进的性能。
英文摘要:
Multimodal representation alignment is pivotal for large language models and robotics. Traditional methods are often hindered by cross-modal information discrepancies and data scarcity, leading to suboptimal alignment spaces that overlook modality-unique features. We propose CodeBind, a framework that optimizes multimodal representation spaces through a modality-shared-specific codebook design. By incrementally aligning target and bridging modalities, CodeBind bypasses the need for fully paired data. Unlike traditional hard alignment, CodeBind decomposes features into shared components for semantic consistency and specific components for modality-unique details. This design utilizes a compositional vector quantization scheme, where a shared codebook bridges modality gaps and modality-specific codebooks mitigate representation bias by preventing dominant modalities from overshadowing others. Validated across nine modalities (text, image, video, audio, depth, thermal, tactile, 3D point cloud, EEG), CodeBind achieves state-of-the-art performance in multimodal classification and retrieval tasks.