多模态流:嵌入空间中语言与视觉的统一流建模
Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces
- Huazhong University of Science and Technology(华中科技大学)
- Beijing Jiaotong University(北京交通大学)
- Horizon Robotics(地平线机器人)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
多模态流提出一种完全连续的生成模型,通过共享分块因果流主干统一语言和视觉,在多个规模上持续预训练,以较少数据达到竞争性能,确立了新的连续多模态范式。
中文摘要 AI 辅助
我们提出了多模态流,一种语言和视觉的完全连续生成模型。大多数统一多模态模型要么将语言和量化图像建模为离散标记,要么将离散语言预测与连续图像生成相结合。前者引入了视觉量化瓶颈。后者需要依赖于模态的目标和采样过程。完全连续的建模避免了这些权衡,并实现了共享的生成过程,但在多模态预训练中仍未得到充分探索。多模态流引入了一种统一的连续架构,将多模态连续表示与共享的分块因果流主干相结合。它将文本块和图像组织为有序的连续超块,保留了文本标记顺序和视觉空间结构。主干通过流匹配在这些超块上学习单一向量场。联合注意力实现了跨模态交互,而特定于模态的前馈网络处理每种模态。模型在训练期间并行预测多个目标块,并在推理时顺序生成超块。我们实例化了MF-1并在多模态数据上进行了预训练。在0.6B、1.2B和1.6B规模下,持续预训练一致地改善了多模态建模。仅使用150B预训练标记,MF-1在GenEval和DPG-Bench上平均得分为82.8,在VQAv2、MMBench和POPE上平均得分为75.3,与在更多数据上训练的统一模型保持竞争力。在匹配的数据、优化和参数预算下,多模态流进一步优于代表性的混合和离散模型。这些结果确立了基于连续块嵌入的流建模作为统一多模态建模的新完全连续范式。相关代码和模型已公开发布于该https URL。
英文摘要
We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with continuous image generation. The former introduces a visual quantization bottleneck. The latter requires modality-dependent objectives and sampling procedures. Fully continuous modeling avoids these trade-offs and enables a shared generative process, but remains underexplored for multimodal pretraining. Multimodal Flow introduces a unified continuous architecture that integrates multimodal continuous representations with a shared chunk-causal flow backbone. It organizes text blocks and images as ordered continuous hyperchunks, preserving textual token order and visual spatial structure. The backbone learns a single vector field over these hyperchunks through Flow Matching. Joint attention enables cross-modal interaction, while modality-specific feed-forward networks process each modality. The model predicts multiple target chunks in parallel during training and generates hyperchunks sequentially at inference. We instantiate MF-1 and pretrain it on multimodal data. Across 0.6B, 1.2B, and 1.6B scales, continued pretraining consistently improves multimodal modeling. With only 150B pretraining tokens, MF-1 achieves an average score of 82.8 across GenEval and DPG-Bench and 75.3 across VQAv2, MMBench, and POPE, remaining competitive with unified models trained on substantially more data. Under matched data, optimization, and parameter budgets, Multimodal Flow further outperforms representative hybrid and discrete models. These results establish continuous chunk-based embedding flow modeling as a new fully continuous paradigm for unified multimodal modeling. The related code and model are publicly released at https://github.com/hustvl/Multimodal-Flow.