MLLM as Retriever: Interactively Learning Multimodal Retrieval for Embodied Agents
Junpeng Yue, Xinrun Xu, Börje F. Karlsson, Zongqing Lu
机构
*
School of Computer Science, Peking University(北京大学计算机科学学院)
;
Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
Jina CLIP: Your CLIP Model Is Also Your Text Retriever
Andreas Koukounas, Georgios Mastrapas, Michael Günther, Bo Wang, Scott Martens, Isabelle Mohr, Saba Sturua, Mohammad Kalim Akram, Joan Fontanals Martínez, Saahil Ognawala, Susana Guzman, Maximilian Werk, Nan Wang, Han Xiao
机构
*
School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)
;
Beijing Key Laboratory of Artificial Intelligence for Education(北京人工智能教育重点实验室)
;
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
Department of Computer Science and Technology, Institute for AI, Tsinghua University(清华大学人工智能研究院计算机科学与技术系)
;
Northeastern University(东北大学)