WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
WeMMU: 通过噪声查询标记增强视觉-语言模型与扩散模型的桥梁
机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知联合实验室,中国科学技术大学) ; ZheJiang University(浙江大学) ; The Hong Kong University of Science and Technology(香港科技大学)
专题命中 VLM训练与架构 :vision-language model(title,abstract);VLM(abstract);multimodal large language model(abstract);分类 cs.CV
AI总结 WeMMU通过噪声查询标记和VAE分支,提升视觉-语言模型与扩散模型的连接效率,缓解泛化崩溃问题,实现稳定持续学习。