FLAT:将图像和文本重采样为1D灵活长度对齐的跨模态标记,用于检索和生成
FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
浏览论文内容
中文总结 AI 辅助
FLAT通过联合优化多模态编码器与生成解码器,将表示学习和生成统一,实现灵活长度的跨模态检索与生成,并在多个基准上达到先进性能。
中文摘要 AI 辅助
传统的多模态表示学习和生成是两个阶段:首先训练一个对比或自监督的视觉编码器,然后训练一个单独的下游生成模型。这种设置将生成性能瓶颈限制在冻结的嵌入上。为弥合这一差距,我们重新审视联合多模态表示学习和生成,以产生可直接被生成解码器使用的线性可插值嵌入。我们提出FLAT(灵活长度对齐的跨模态表示),一个表示预训练框架,它联合优化共享的多模态编码器以及下游的文本到图像(T2I)和图像到文本(I2T)解码器。通过将对比对齐与双向跨模态生成目标相结合,FLAT确保其表示既作为判别性语义描述符,又作为生成条件。在架构上,FLAT将视觉和文本输入映射到统一的连续1D序列空间,对前缀K标记应用嵌套dropout以实现动态输出长度。单个预训练阶段允许FLAT在可变前缀K下执行跨模态检索和生成,在T2I GenEval上达到71.1的分数。任务特定的微调使模型性能与最先进的基线对齐:T2I生成上GenEval为83.1;MS-COCO图像描述上BLEU-4为40.5,CIDEr为138.6;MS-COCO上Recall@5为86.8(I2T)/75.8(T2I),Flickr30K上为98.3(I2T)/93.6(T2I)。最后,定性评估表明FLAT表示原生支持线性插值、潜在空间算术和零样本组合检索。
英文摘要
Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream generative model. This setup bottlenecks generative performance behind frozen embeddings. To bridge this gap, we revisit joint multimodal representation learning and generation to produce linearly interpolatable embeddings that are directly consumable by generative decoders. We present FLAT (Flexible-Length Aligned Transmodal representations), a representation pre-training framework that jointly optimizes a shared multimodal encoder alongside downstream text-to-image (T2I) and image-to-text (I2T) decoders. By combining contrastive alignment with bidirectional cross-modal generative objectives, FLAT ensures its representations function as both discriminative semantic descriptors and generative conditions. Architecturally, FLAT maps visual and textual inputs into a unified continuous 1D sequence space, applying nested dropout over prefix-K tokens to enable dynamic output lengths. A single pre-training stage allows FLAT to perform cross-modal retrieval and generation across variable prefix K, achieving a T2I GenEval score of 71.1. Task-specific fine-tuning aligns model performance with state-of-the-art baselines: 83.1 GenEval on T2I generation; 40.5 BLEU-4 and 138.6 CIDEr on MS-COCO image captioning; and Recall@5 scores of 86.8 (I2T) / 75.8 (T2I) on MS-COCO alongside 98.3 (I2T) / 93.6 (T2I) on Flickr30K. Finally, qualitative evaluations demonstrate that FLAT representations natively support linear interpolation, latent space arithmetic, and zero-shot composed retrieval.
发表机构
- Meta AI
机构由 AI 辅助整理,请以论文原文为准。