Decoupled Vision-Language System for Multimodal Understanding and Generation
用于多模态理解与生成的解耦式视觉-语言系统
机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; University of Chinese Academy of Sciences(中国科学院大学) ; Peng Cheng Laboratory(鹏城实验室) ; Li Auto(理想汽车)
专题命中 文生图 :image generation(abstract);text-to-image(abstract);分类 cs.CV
AI总结 本研究提出多模态大语言模型Libra的解耦式视觉-语言架构,通过开关注意力与FFN模块实现自模态与跨模态解耦,在理解与生成任务上均取得优异性能。