Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
机构 * University of California, San Diego(加州大学圣迭戈分校) ; Adobe Research(Adobe研究)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Preprint
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * University of California, San Diego(加州大学圣迭戈分校) ; Adobe Research(Adobe研究)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
Comments Preprint
机构 * AIDAS Laboratory(AIDAS实验室) ; IPAI ; ECE(电子工程系) ; Seoul National University(首尔国立大学)
专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI
Comments Accepted in NeurIPS 2025
机构 * Peter Munk Cardiac Centre, University Health Network (UHN)(彼得·默克心脏中心,大学健康网络) ; Department of Medical Biophysics, UofT(医学生物物理学系) ; Ted Rogers Centre for Heart Research, UHN(泰德·罗杰斯心脏病研究中心,大学健康网络) ; Department of Computer Science, University of Toronto (UofT)(计算机科学系,多伦多大学) ; Toronto General Hospital Research Institute, UHN(多伦多总医院研究 institute) ; Department of Medical Imaging, UofT(医学影像学系) ; Vector Institute, Toronto(向量研究所)
专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV
Comments MICCAI 2025
Journal ref Medical Image Computing and Computer Assisted Intervention - MICCAI 2025. MICCAI 2025. Lecture Notes in Computer Science, vol 15964. Springer, Cham
专题命中 图文多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI
Comments First Peer Reviewed Review Paper for Object Detection with Vision-Language Models (VLMs)
Journal ref Information Fusion, 2025
机构 * School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院) ; College of Artificial Intelligence, Xi’an Jiaotong University(西安交通大学人工智能学院) ; School of Computer Science and Technology and Ministry of Education Key Lab For Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院和教育部智能网络与网络安全重点实验室) ; School of Mathematics and Statistics and Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学数学与统计学院和教育部智能网络与网络安全重点实验室) ; Pazhou Laboratory (Huangpu), Guangzhou, Guangdong, China(琶洲实验室(黄埔),广州,广东,中国)
专题命中 图文多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV
机构 * University of Manchester(曼彻斯特大学) ; South China University of Technology(华南理工大学)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments Accept by EMNLP2025
机构 * Zhuoning Xu 1(Xu Zhuoning 1) ; Xinyan Liu 1(Liu Xinyan 1)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract)
机构 * VUNO Inc.(VUNO公司) ; KAIST(韩国科学技术院)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI
Comments 38 pages, 17 figures, preprint
机构 * Birla Institute of Technology and Science Pilani, K.K Birla Goa Campus(比拉理工学院和科学学院,K.K比拉果阿校区)
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV
Journal ref Proc. European Conference on Mobile Robots (ECMR), 2025, pp. 1-6