arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3460 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3460 篇

2507.23331 2025-08-01 cs.CV 83%

Contrastive Learning-Driven Traffic Sign Perception: Multi-Modal Fusion of Text and Vision

Qiang Lu, Waikit Xiu, Xiying Li, Shenyu Hu, Shengbo Sun

机构 * School of Intelligent Systems Engineering, Sun Yat-sen University(中山大学智能系统工程学院) Guangdong Provincial Key Laboratory of Intelligent Transportation System(广东省智能交通系统重点实验室)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 11pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21917 2025-07-30 cs.CV 83%

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

Nicola Fanelli, Gennaro Vessio, Giovanna Castellano

机构 * Department of Computer Science University of Bari Aldo Moro(计算机科学系巴里大学Aldo Moro)

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20326 2025-07-29 cs.LG cs.AI 83%

MIPS: a Multimodal Infinite Polymer Sequence Pre-training Framework for Polymer Property Prediction

Jiaxi Wang, Yaosen Min, Xun Zhu, Miao Li, Ji Wu

机构 * Department of Electronic Engineering, Tsinghua University Beijing China Zhongguancun Institute of Artificial Intelligence Beijing China Department of Electronic Engineering \& College of AI, Tsinghua University Beijing National Research Center for Information Science Department of Electronic Engineering, Tsinghua University Zhongguancun Institute of Artificial Intelligence

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 14 pages, 8 figures, accepted by ACM Multimedia 2025 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01334 2025-07-22 cs.IR cs.CV 83%

Composed Multi-modal Retrieval: A Survey of Approaches and Applications

Kun Zhang, Jingyu Li, Zhe Li, Jingjing Zhang, Fan Li, Yandong Liu, Rui Yan, Zihang Jiang, Nan Chen, Lei Zhang, Yongdong Zhang, Zhendong Mao, S. Kevin Zhou

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22589 2025-07-15 cs.CV 83%

LIGHT: Multi-Modal Text Linking on Historical Maps

Yijun Lin, Rhett Olson, Junhan Wu, Yao-Yi Chiang, Jerod Weinman

机构 * University of Minnesota(明尼苏达大学) Grinnell College(格林内尔学院)

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at ICDAR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21538 2025-06-27 cs.CV cs.IR cs.LG 83%

Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval

Hani Alomari, Anushka Sivakumar, Andrew Zhang, Chris Thomas

机构 * Virginia Tech(弗吉尼亚理工大学)

专题命中 跨模态检索 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted at the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17782 2025-06-24 cs.IR cs.AI 83%

Expanding Relevance Judgments for Medical Case-based Retrieval Task with Multimodal LLMs

Catarina Pires, Sérgio Nunes, Luís Filipe Teixeira

机构 * Faculty of Engineering, University of Porto(葡萄牙波尔图大学工程学院)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments To appear at the Third Workshop on Large Language Models for Evaluation in Information Retrieval (LLM4Eval 2025), co-located with SIGIR 2025. 9 pages, 2 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.05929 2025-06-19 cs.LG cs.AI 83%

M3-JEPA: Multimodal Alignment via Multi-gate MoE based on the Joint-Embedding Predictive Architecture

Hongyang Lei, Xiaolong Cheng, Qi Qin, Dan Wang, Kun Fan, Huazhen Huang, Qingqing Gu, Yetao Wu, Zhonglin Jiang, Yong Chen, Luo Ji

机构 * Geely AI Lab(Geely人工智能实验室) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Peking University(北京大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 16 pages, 5 figures. ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11674 2025-06-16 cs.CV 83%

Cross-Modal Clustering-Guided Negative Sampling for Self-Supervised Joint Learning from Medical Images and Reports

Libin Lan, Hongxing Li, Zunhui Xia, Juan Zhou, Xiaofei Zhu, Yongmei Li, Yudong Zhang, Xin Luo

机构 * College of Computer Science and Engineering, Chongqing University of Technology(重庆理工大学计算机科学与工程学院)

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments This work has been submitted to the IEEE TMI for possible publication. Our code is available at https://github.com/violet-42/CM-CGNS

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11184 2025-06-04 cs.CL cs.AI cs.CV cs.MM 83%

Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs

Wenxuan Wang, Xiaoyuan Liu, Kuiyi Gao, Jen-tse Huang, Youliang Yuan, Pinjia He, Shuai Wang, Zhaopeng Tu

机构 * Renmin University of China(中国人民大学) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Chinese University of Hong Kong(香港中文大学) Johns Hopkins University(约翰霍普金斯大学) Hong Kong University of Science and Technology(香港科技大学) Tencent(腾讯)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02020 2025-06-04 cs.CV cs.LG 83%

Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying

Youze Xue, Dian Li, Gang Liu

机构 * Tencent(腾讯)

专题命中 跨模态检索 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24361 2025-06-02 cs.CV 83%

Revisiting Cross-Modal Knowledge Distillation: A Disentanglement Approach for RGBD Semantic Segmentation

Roger Ferrod, Cássio F. Dantas, Luigi Di Caro, Dino Ienco

机构 * University of Turin(都灵大学) INRAE, UMR TETIS, Univ. Montpellier(法国蒙彼利埃大学、INRAE、UMR TETIS) EVERGREEN, Univ. Montpellier, Inria(法国蒙彼利埃大学、EVERGREEN、Inria)

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19707 2025-05-27 cs.CV cs.IR 83%

MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval

Rong-Cheng Tu, Zhao Jin, Jingyi Liao, Xiao Luo, Yingjie Wang, Li Shen, Dacheng Tao

机构 * College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算机与数据科学学院) Department of Computer Science, University of California, Los Angeles, USA(加州大学洛杉矶分校计算机科学系) Sun Yat-sen University Shenzhen Campus, School of Cyber Science and Technology, Shenzhen, China(中山大学深圳校区信息科学与技术学院)

专题命中 跨模态检索 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02087 2025-05-06 cs.AI 83%

Retrieval-augmented in-context learning for multimodal large language models in disease classification

Zaifu Zhan, Shuang Zhou, Xiaoshan Zhou, Yongkang Xiao, Jun Wang, Jiawen Deng, He Zhu, Yu Hou, Rui Zhang

专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments 17 Pages, 1 figure, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15344 2025-05-06 cs.SD eess.AS 83%

Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions

Yifei Xin, Yuexian Zou

专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 eess.AS

Comments Accepted by Interspeech2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21044 2025-05-01 cs.CR cs.AI 83%

AGATE: Stealthy Black-box Watermarking for Multimodal Model Copyright Protection

Jianbo Gao, Keke Gai, Jing Yu, Liehuang Zhu, Qi Wu

机构 * Beijing Institute of Technology(北京理工大学) School of Information Engineering, Minzu University of China(民族大学信息工程学院) University of Adelaide(阿德莱德大学)

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02698 2025-04-29 cs.LG cs.AI q-bio.QM 83%

SCMPPI: Supervised Contrastive Multimodal Framework for Predicting Protein-Protein Interactions

Shengrui XU, Tianchi Lu, Zikun Wang, Jixiu Zhai

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 20 pages,9 figures,conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16576 2025-04-24 cs.IR cs.AI 83%

MMHCL: Multi-Modal Hypergraph Contrastive Learning for Recommendation

Xu Guo, Tong Zhang, Fuyun Wang, Xudong Wang, Xiaoya Zhang, Xin Liu, Zhen Cui

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(计算机科学与工程学院,南京理工大学) Shituoyun (Nanjing) Technology Co., Ltd(石墨云(南京)科技有限公司) School of Artificial Intelligence, Beijing Normal University(人工智能学院,北京师范大学)

专题命中 跨模态检索 :multi-modal(title,abstract);multimodal(abstract);分类 cs.AI

Comments 23 pages, 8 figures. This manuscript is currently under major revision for ACM Transactions on Multimedia Computing, Communications, and Applications (ACM TOMM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10074 2025-04-22 cs.AI 83%

MMKB-RAG: A Multi-Modal Knowledge-Based Retrieval-Augmented Generation Framework

Zihan Ling, Zhiyao Guo, Yixuan Huang, Yi An, Shuai Xiao, Jinsong Lan, Xiaoyong Zhu, Bo Zheng

机构 * Peking University(北京大学) Alibaba Group(阿里巴巴集团)

专题命中 跨模态检索 :multi-modal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02264 2025-04-04 cs.CV 83%

MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving Perception

Wenzhuo Liu, Wenshuo Wang, Yicheng Qiao, Qiannan Guo, Jiayin Zhu, Pengfei Li, Zilong Chen, Huiming Yang, Zhiwei Li, Lening Wang, Tiao Tan, Huaping Liu

专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01476 2025-04-03 cs.CV 83%

Enhanced Cross-modal 3D Retrieval via Tri-modal Reconstruction

Junlong Ren, Hao Wang

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14977 2025-04-03 cs.AI eess.IV 83%

Trustworthy Enhanced Multi-view Multi-modal Alzheimer's Disease Prediction with Brain-wide Imaging Transcriptomics Data

Shan Cong, Zhoujie Fan, Hongwei Liu, Yinghan Zhang, Xin Wang, Haoran Luo, Xiaohui Yao

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16855 2025-04-02 cs.CL cs.IR 83%

GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Xin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li, Ziqi Dai, Dingkun Long, Pengjun Xie, Meishan Zhang, Wenjie Li, Min Zhang

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments Accepted to CVPR 2025, models at https://huggingface.co/Alibaba-NLP/gme-Qwen2-VL-2B-Instruct

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12287 2025-03-24 cs.CL 83%

CUE-M: Contextual Understanding and Enhanced Search with Multimodal Large Language Model

Dongyoung Go, Taesun Whang, Chanhee Lee, Hwa-Yeon Kim, Sunghoon Park, Seunghwan Ji, Jinho Kim, Dongchan Kim, Young-Bum Kim

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments Preprint. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12914 2025-03-18 cs.CV 83%

Efficient Multimodal 3D Object Detector via Instance-Level Contrastive Distillation

Zhuoqun Su, Huimin Lu, Shuaifeng Jiao, Junhao Xiao, Yaonan Wang, Xieyuanli Chen

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06277 2025-03-18 cs.CV 83%

STiL: Semi-supervised Tabular-Image Learning for Comprehensive Task-Relevant Information Exploration in Multimodal Classification

Siyi Du, Xinzhe Luo, Declan P. O'Regan, Chen Qin

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 16 pages (including 5 pages of supplementary materials), accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18363 2025-03-12 cs.CV 83%

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Qing Jiang, Gen Luo, Yuqin Yang, Yuda Xiong, Yihao Chen, Zhaoyang Zeng, Tianhe Ren, Lei Zhang

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments 35 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15268 2025-02-07 cs.CL 83%

Fact-Aware Multimodal Retrieval Augmentation for Accurate Medical Radiology Report Generation

Liwen Sun, James Zhao, Megan Han, Chenyan Xiong

专题命中 跨模态检索 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CL

Comments NAACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07157 2025-01-14 cs.AI 83%

CureGraph: Contrastive Multi-Modal Graph Representation Learning for Urban Living Circle Health Profiling and Prediction

Jinlin Li, Xiao Zhou

专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01720 2024-12-03 cs.CV 83%

LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant

Yikun Liu, Pingan Chen, Jiayin Cai, Xiaolong Jiang, Yao Hu, Jiangchao Yao, Yanfeng Wang, Weidi Xie

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏