Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
机构 * Google DeepMind(谷歌DeepMind) ; University of Central Florida(中央佛罗里达大学)
专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Google DeepMind(谷歌DeepMind) ; University of Central Florida(中央佛罗里达大学)
专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI
专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI
Comments 10 pages, 10 figures, CoLM 2025
专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL
Comments 33 pages, 11 figures, 7 tables
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
机构 * Intelligent Game and Decision Lab(智能游戏与决策实验室)
专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV
Comments IEEE Submission
机构 * VLM Safety LAB, MODULABS(视觉语言模型安全实验室,MODULABS) ; ETRI(电子技术研究院) ; KAIST(韩国科学技术院)
专题命中 图文多模态 :multimodal(abstract,comments);image-text(abstract);分类 cs.CV、cs.AI
Comments Accepted to Safe and Trustworthy Multimodal AI Systems(SafeMM-AI) Workshop at ICCV2025, Non-archival track
机构 * The Hong Kong Polytechnic University(香港理工大学) ; Tsinghua University(清华大学) ; InspireOmni AI ; Alibaba Group(阿里巴巴集团) ; Case Western Reserve University(凯斯西储大学)
专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL
Comments Accepted at NeurIPS 2025
机构 * CSE Department, HKUST(香港科技大学计算机科学与工程系) ; Tencent AI Seattle Lab(腾讯AI西雅图实验室) ; University of Edinburgh(爱丁堡大学) ; NVIDIA AI Technology Center (NVAITC), NVIDIA, Santa Clara, USA(英伟达圣克拉拉人工智能技术中心)
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.CL
Comments Accepted as a spotlight at NeurIPS 2025
机构 * Department of Computer Vision, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(计算机视觉系,Mohamed bin Zayed人工智能大学)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
机构 * Florida State University(佛罗里达州立大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 12 pages, 5 figures
机构 * School of Electronic Engineering(电子工程学院) ; Xidian University(西安电子科技大学) ; School of Computer Science(计算机科学学院)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
机构 * NAVER AI Lab(NAVER AI实验室)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
Comments Code: https://github.com/naver-ai/prolip HuggingFace Hub: https://huggingface.co/collections/SanghyukChun/prolip-6712595dfc87fd8597350291 33 pages, 4.5 MB; LongProLIP paper: arXiv:2503.08048; Multiplicity paper for more background: arxiv.org:2505.19614; v4: fix typos
机构 * Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) ; Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州先进研究所) ; Tsinghua University(清华大学)
专题命中 图文多模态 :image-text(abstract)
Comments 13 pages