Vision-Centric Activation and Coordination for Multimodal Large Language Models
机构 * MoE Key Lab of Artificial Intelligence, Shanghai Jiao Tong University(人工智能MoE实验室,上海交通大学) ; Ant Group(蚂蚁集团) ; Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo(宁波数字孪生研究所,东部技术研究所,宁波)
专题命中 VLM训练与架构 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV