arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3475 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3475 篇

2510.18303 2025-10-22 cs.CV 79%

Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language Models

Lehan Wang, Yi Qin, Honglong Yang, Xiaomeng Li

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15547 2025-10-20 cs.AI cs.ET cs.LG cs.SY eess.SP eess.SY 79%

Hypergraph Contrastive Sensor Fusion for Multimodal Fault Diagnosis in Induction Motors

Usman Ali, Ali Zia, Waqas Ali, Umer Ramzan, Abdul Rehman, Muhammad Tayyab Chaudhry, Wei Xiang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Submitted to IEEE Sensors Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09585 2025-10-20 cs.CV 79%

Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation

Jitesh Jain, Zhengyuan Yang, Humphrey Shi, Jianfeng Gao, Jianwei Yang

机构 * Microsoft Research, Redmond(微软研究院(红mond)) Meta Superintelligence Labs(Meta超智能实验室)

专题命中 跨模态检索 :multimodal(title);MLLM(abstract);分类 cs.CV

Comments Project Page: https://praeclarumjj3.github.io/visper_lm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14136 2025-10-17 cs.AI 79%

A Multimodal Approach to Heritage Preservation in the Context of Climate Change

David Roqui, Adèle Cormier, nistor Grozavu, Ann Bourges

机构 * ETIS laboratory, Cergy university(Cergy大学ETIS实验室) Research and Restoration Center for the Museums of France (C2RMF)(法国博物馆研究与修复中心(C2RMF))

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12801 2025-10-15 cs.CV cs.IR 79%

DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search

Kartik Narayan, Yang Xu, Tian Cao, Kavya Nerella, Vishal M. Patel, Navid Shiee, Peter Grasch, Chao Jia, Yinfei Yang, Zhe Gan

机构 * Johns Hopkins University(约翰霍普金斯大学) Apple(苹果公司)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10828 2025-10-14 cs.IR cs.AI 79%

VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering

Zhenghan Tai, Hanwei Wu, Qingchen Hu, Jijun Chi, Hailin He, Lei Ding, Tung Sum Thomas Kwok, Bohuai Xiao, Yuchen Hua, Suyuchen Wang, Peng Lu, Muzhi Li, Yihong Wu, Liheng Ma, Jerry Huang, Jiayi Zhang, Gonghao Zhang, Chaolong Jiang, Jingrui Tian, Sicheng Lyu, Zeyu Li, Boyu Han, Fengran Mo, Xinyue Yu, Yufei Cui, Ling Zhou, Xinyu Wang

机构 * University of Toronto(多伦多大学) McMaster University(麦马斯特大学) McGill University(麦吉尔大学) University of Manitoba(曼尼托巴大学) University of California, Los Angeles(加州大学洛杉矶分校) University of Montreal(蒙特利尔大学) Mila CUHK(香港中文大学) HKUST(GZ)(香港理工大学(广州)) Nanyang Technological University(南洋理工大学) Stanford University(斯坦福大学) CG Matrix Technology Limited(CG矩阵科技有限公司)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04379 2025-10-13 cs.CV cs.LG 79%

VisionTS++: Cross-Modal Time Series Foundation Model with Continual Pre-trained Vision Backbones

Lefei Shen, Mouxiang Chen, Xu Liu, Han Fu, Xiaoxue Ren, Jianling Sun, Zhuo Li, Chenghao Liu

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学) State Street Technology (Zhejiang) Ltd.(State Street Technology(浙江)有限公司) Salesforce Research Asia(Salesforce亚洲研究)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03268 2025-10-09 cs.LG cs.AI 79%

Decipher the Modality Gap in Multimodal Contrastive Learning: From Convergent Representations to Pairwise Alignment

Lingjie Yi, Raphael Douady, Chao Chen

机构 * Stony Brook University(史坦尼·布鲁克大学) University Paris 1 Pantheon-Sorbonne(巴黎第一大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11465 2025-10-07 cs.CL cs.LG 79%

CEMTM: Contextual Embedding-based Multimodal Topic Modeling

Amirhossein Abaskohi, Raymond Li, Chuyuan Li, Shafiq Joty, Giuseppe Carenini

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00579 2025-10-06 cs.MM cs.IR 79%

MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning

Ziyu Gong, Chengcheng Mai, Yihua Huang

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.MM

Comments Comments: Update Title, Author, Abstract, etc

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02580 2025-10-06 cs.AI 79%

V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving

Xuewen Luo, Fengze Yang, Fan Ding, Xiangbo Gao, Shuo Xing, Yang Zhou, Zhengzhong Tu, Chenxi Liu

机构 * University of Utah(犹他大学) Monash University(莫纳什大学) Texas A&M University(德克萨斯农工大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06461 2025-10-06 cs.CV 79%

Ranked from Within: Ranking Large Multimodal Models Without Labels

Weijie Tu, Weijian Deng, Dylan Campbell, Yu Yao, Jiyang Zheng, Tom Gedeon, Tongliang Liu

机构 * Australian National University Sydney AI Centre, The University of Sydney Curtin University University of \'OBuda

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments ICML 2025 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25711 2025-10-01 cs.CV 79%

ProbMed: A Probabilistic Framework for Medical Multimodal Binding

Yuan Gao, Sangwook Kim, Jianzhong You, Chris McIntosh

机构 * Peter Munk Cardiac Centre(彼得·默克心脏中心) Ted Rogers Centre for Heart Research(泰德·罗杰斯心脏病研究中心) University Health Network(大学健康网络) Joint Department of Medical Imaging(联合医学影像部门) University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26378 2025-10-01 cs.IR cs.CV 79%

MR$^2$-Bench: Going Beyond Matching to Reasoning in Multimodal Retrieval

Junjie Zhou, Ze Liu, Lei Xiong, Jin-Ge Yao, Yueze Wang, Shitao Xiao, Fenfen Lin, Miguel Hu Chen, Zhicheng Dou, Siqi Bao, Defu Lian, Yongping Xiong, Zheng Liu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21151 2025-09-26 cs.CL cs.IR 79%

Retrieval over Classification: Integrating Relation Semantics for Multimodal Relation Extraction

Lei Hei, Tingjing Liao, Yingxin Pei, Yiyang Qi, Jiaqi Wang, Ruiting Li, Feiliang Ren

机构 * School of Computer Science and Engineering, Northeastern University, Shenyang 110819, China(计算机科学与工程学院,东北大学,沈阳110819,中国)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20501 2025-09-26 cs.LG cs.CV 79%

Beyond Visual Similarity: Rule-Guided Multimodal Clustering with explicit domain rules

Kishor Datta Gupta, Mohd Ariful Haque, Marufa Kamal, Ahmed Rafi Hasan, Md. Mahfuzur Rahman, Roy George

机构 * Clark Atlanta University(克拉克阿特兰大学) BRAC University(布拉克大学) United International University(国际联合大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15141 2025-09-26 cs.AI cs.LG physics.chem-ph 79%

Text-Augmented Multimodal LLMs for Chemical Reaction Condition Recommendation

Yu Zhang, Ruijie Yu, Kaipeng Zeng, Ding Li, Feng Zhu, Xiaokang Yang, Yaohui Jin, Yanyan Xu

机构 * Institute for Clarity in Documentation(清晰文档研究所) Inria Paris-Rocquencourt(巴黎-罗克琴克研究所) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕勒尔研究实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19965 2025-09-25 cs.CV 79%

SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding

Phyo Thet Yee, Dimitrios Kollias, Sudeepta Mishra, Abhinav Dhall

机构 * IIT Ropar(印度IIT罗帕尔) Queen Mary University of London(伦敦女王玛丽大学) Monash University(墨尔本大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at WACV 2026, project page : https://novicemm.github.io/synchrorama

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19203 2025-09-24 cs.CV 79%

Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions

Ioanna Ntinou, Alexandros Xenos, Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

机构 * Queen Mary University of London(伦敦女王大学) Samsung AI Centre(三星人工智能中心) Technical University of Iași(伊阿苏技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18284 2025-09-24 cs.CV 79%

Learning Contrastive Multimodal Fusion with Improved Modality Dropout for Disease Detection and Prediction

Yi Gu, Kuniaki Saito, Jiaxin Ma

机构 * OMRON SINIC X Corporation(OMRON SINIC X公司) Nara Institute of Science and Technology(名取科学技術大學院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17762 2025-09-23 cs.CV 79%

Neural-MMGS: Multi-modal Neural Gaussian Splats for Large-Scale Scene Reconstruction

Sitian Shen, Georgi Pramatarov, Yifu Tao, Daniele De Martini

机构 * Mobile Robotics Group (MRG), Oxford Robotics Institute, Department of Engineering Science, University of Oxford, UK(移动机器人组(MRG),牛津机器人研究所,工程科学系,牛津大学,英国)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15269 2025-09-23 cs.CV cs.LG 79%

Test-Time Multimodal Backdoor Detection by Contrastive Prompting

Yuwei Niu, Shuo He, Qi Wei, Zongyu Wu, Feng Liu, Lei Feng

机构 * Chongqing University(重庆大学) Nanyang Technological University(南洋理工大学) Penn State University(宾夕法尼亚州立大学) University of Melbourne(墨尔本大学) Southeast University(东南大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to ICML2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15211 2025-09-19 cs.CL 79%

What's the Best Way to Retrieve Slides? A Comparative Study of Multimodal, Caption-Based, and Hybrid Retrieval Techniques

Petros Stylianos Giouroukis, Dimitris Dimitriadis, Dimitrios Papadopoulos, Zhenwen Shao, Grigorios Tsoumakas

机构 * Aristotle University of Thessaloniki(亚里士多德大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12590 2025-09-19 cs.CV cs.LG 79%

Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering

Neale Ratzlaff, Matthew Lyle Olson, Musashi Hinck, Estelle Aflalo, Shao-Yen Tseng, Vasudev Lal, Phillip Howard

机构 * Oracle Intel Labs(英特尔实验室) Thoughtworks

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages, 6 Figures, 8 Tables. arXiv admin note: text overlap with arXiv:2410.13976

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13474 2025-09-18 cs.CV 79%

Semantic-Enhanced Cross-Modal Place Recognition for Robust Robot Localization

Yujia Lin, Nicholas Evans

机构 * Dali University(大理大学) Bandırma Onyedi Eylül University(巴尔迪马十一点大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10282 2025-09-15 cs.CV cs.LG 79%

MCL-AD: Multimodal Collaboration Learning for Zero-Shot 3D Anomaly Detection

Gang Li, Tianjiao Chen, Mingle Zhou, Min Li, Delong Han, Jin Wan

机构 * Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan)(计算能力网络与信息安全部教育部重点实验室,山东计算机科学中心(济南国家超级计算机中心)) Qilu University of Technology (Shandong Academy of Sciences)(齐鲁工业大学(山东科学院)) Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science(山东省计算能力互联网与服务计算重点实验室,山东省计算机科学基础研究中心)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Page 14, 5 pictures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07666 2025-09-10 cs.CL cs.IR 79%

MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval

Xixi Wu, Yanchao Tan, Nan Hou, Ruiyang Zhang, Hong Cheng

机构 * The Chinese University of Hong Kong(香港中文大学) Fuzhou University(福州大学) University of Macau(澳门大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

Comments EMNLP Main 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02017 2025-09-03 cs.IR cs.AI 79%

Empowering Large Language Model for Sequential Recommendation via Multimodal Embeddings and Semantic IDs

Yuhao Wang, Junwei Pan, Xinhang Li, Maolin Wang, Yuan Wang, Yue Liu, Dapeng Liu, Jie Jiang, Xiangyu Zhao

机构 * City University of Hong Kong(香港城市大学) Tencent Inc.(腾讯公司) Tsinghua University(清华大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments CIKM 2025 Full Research Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00751 2025-09-03 cs.CV 79%

EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions

Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science, VNU-HCM(越南国家大学科学学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21595 2025-09-03 cs.CV 79%

PS-ReID: Advancing Person Re-Identification and Precise Segmentation with Multimodal Retrieval

Jincheng Yan, Yun Wang, Xiaoyan Luo, Yu-Wing Tai

机构 * School of Astronautics, Beihang University(北京航空航天大学航天学院) Department of Computer Science, Dartmouth College(达特茅斯学院计算机科学系)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏