Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
机构 * aThe State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, China [1ex] bSchool of Artificial Intelligence, University of Chinese Academy of Sciences, China [1ex] cSchool of Artificial Intelligence, Beijing University of Posts ; Telecommunications, China [1ex] dKey Laboratory of Computing Power Network ; Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences), China [1ex] e Shandong Provincial Key Laboratory of Computing Power Internet ; Service Computing, Shandong Fundamental Research Center for Computer Science, China
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments 27 pages, 11 figures. Accepted to Information Fusion. Final journal version: volume 126 (Part B), February 2026
Journal ref Information Fusion, 126 (Part B), February 2026, 103652