Omni-Embed-Nemotron: A Unified Multimodal Retrieval Model for Text, Image, Audio, and Video
机构 * NVIDIA
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * NVIDIA
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL
Comments EMNLP 2025
机构 * Rochester Institute of Technology(罗切斯特技术研究所)
专题命中 跨模态检索 :multimodal(abstract);multimodal foundation model(abstract)
专题命中 跨模态检索 :multimodal(abstract);分类 cs.MM
Comments 28 Pages, 17 Figures
机构 * University of Alberta(阿尔伯塔大学) ; Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所) ; Charles University(查尔斯大学) ; EquiLibre Technologies, Inc.(EquiLibre技术公司) ; Allen Institute for AI(艾伦人工智能研究所) ; Sony AI(索尼人工智能)
专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI
专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV
机构 * CISPA Helmholtz Center for Information Security(CISPA 欧洲信息安全研究中心)
专题命中 跨模态检索 :cross-modal(abstract)