arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3460 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3460 篇

2511.21707 2025-12-01 cs.NI cs.AI 83%

Sensing and Understanding the World over Air: A Large Multimodal Model for Mobile Networks

通过空气感知和理解世界:为移动网络设计的大型多模态模型

Zhuoran Duan, Yuhao Wei, Guoshun Nan, Zijun Wang, Yan Yan, Lihua Xiong, Yuhan Ran, Ji Zhang, Jian Li, Qimei Cui, Xiaofeng Tao, Tony Q. S. Quek

机构 * National Engineering Research Center for Mobile Network Technologies, Beijing University of Posts and Telecommunications (BUPT), Beijing(中国移动网络技术国家工程研究中心,北京邮电大学) Beiyou Shenzhen Institute(北邮深圳研究所) School of Cyber Security, University of Chinese Academy of Sciences (UCAS)(中国科学院大学网络安全学院) China Telecom Co., Ltd.(中国电信股份有限公司) Singapore University of Technology and Design (SUTD)(新加坡科技设计大学)

专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

AI总结 本文提出了一种无线原生多模态模型,利用无线信号进行对比学习,验证了其在无线网络中的应用潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19278 2025-11-27 cs.CV 83%

ReMatch: Boosting Representation through Matching for Multimodal Retrieval

ReMatch: 通过匹配提升表示以实现多模态检索

Qianying Liu, Xiao Liang, Zhiqiang Zhang, Zhongfei Qing, Fengfan Zhou, Yibo Chen, Xu Tang, Yao Hu, Paul Henderson

机构 * University of Glasgow(格拉斯哥大学) Xiaohongshu Inc.(小红书公司) Huazhong University of Science and Technology(华中科技大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 ReMatch通过生成匹配阶段提升多模态检索的表示能力,利用端到端训练和细粒度嵌入生成,在MMEB基准上取得新突破。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19834 2025-11-26 cs.CV 83%

Large Language Model Aided Birt-Hogg-Dube Syndrome Diagnosis with Multimodal Retrieval-Augmented Generation

基于多模态检索增强生成的大型语言模型辅助Birt-Hogg-Dube综合征诊断

Haoqing Li, Jun Shi, Xianmeng Chen, Qiwei Jia, Rui Wang, Wei Wei, Hong An, Xiaowen Hu

机构 * School of Computer Science and Technology(计算机科学与技术学院) Department of Pulmonary and Critical Care Medicine(呼吸与危重症医学科) Center for Diagnosis and Management of Rare Diseases(罕见病诊断与管理中心) the First Affiliated Hospital of USTC(USTC第一附属医院) Division of Life Sciences and Medicine(生命科学与医学系) USTC WanNan Medical College(皖南医学院)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本研究提出BHD-RAG框架,通过整合多模态检索增强生成与领域专业知识,提升Birt-Hogg-Dube综合征的CT影像诊断准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16654 2025-11-25 cs.CL 83%

Comparison of Text-Based and Image-Based Retrieval in Multimodal Retrieval Augmented Generation Large Language Model Systems

多模态检索增强生成大语言模型系统中基于文本和基于图像的检索比较

Elias Lumer, Alex Cardenas, Matt Melich, Myles Mason, Sara Dieter, Vamse Kumar Subbiah, Pradeep Honaganahalli Basavaraju, Roberto Hernandez

机构 * PricewaterhouseCoopers U.S.(普华永道美国公司)

专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

AI总结 本文比较了多模态RAG系统中基于文本和基于图像的检索方法,发现直接多模态嵌入检索在性能和准确性上优于基于LLM总结的方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01588 2025-11-24 cs.LG cs.CV 83%

Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization

探索更多,学习更好:基于互信息最小化的并行 MLLM 嵌入

Zhicheng Wang, Chen Ju, Xu Chen, Shuai Xiao, Jinsong Lan, Xiaoyong Zhu, Ying Chen, Zhiguo Cao

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团) Huazhong University of Science and Technology(华中科技大学)

专题命中 跨模态检索 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文提出基于互信息最小化的并行解耦框架,通过多条并行路径提升多模态嵌入质量,实现高效且鲁棒的嵌入空间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11216 2025-11-17 cs.CV 83%

Positional Bias in Multimodal Embedding Models: Do They Favor the Beginning, the Middle, or the End?

Kebin Wu, Fatima Albreiki

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments accepted to AAAI 2026 main track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10997 2025-11-17 cs.CV cs.LG 83%

PROMISE: Prompt-Attentive Hierarchical Contrastive Learning for Robust Cross-Modal Representation with Missing Modalities

Jiajun Chen, Sai Cheng, Yutao Yuan, Yirui Zhang, Haitao Yuan, Peng Peng, Yi Zhong

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted by AAAI'2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10641 2025-11-07 cs.LG cs.AI 83%

Test-Time Warmup for Multimodal Large Language Models

Nikita Rajaneesh, Thomas Zollo, Richard Zemel

机构 * Columbia University(哥伦比亚大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17784 2025-10-31 cs.LG cs.AI 83%

Revealing Multimodal Causality with Large Language Models

Jin Li, Shoujin Wang, Qi Zhang, Feng Liu, Tongliang Liu, Longbing Cao, Shui Yu, Fang Chen

机构 * University of Technology Sydney(技术大学悉尼大学) Tongji University(同济大学) University of Melbourne(墨尔本大学) University of Sydney(悉尼大学) Macquarie University(麦考瑞大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06888 2025-10-09 cs.IR cs.AI 83%

M3Retrieve: Benchmarking Multimodal Retrieval for Medicine

Arkadeep Acharya, Akash Ghosh, Pradeepika Verma, Kitsuchart Pasupa, Sriparna Saha, Priti Singh

机构 * Indian Institute of Technology Patna(印度理工学院帕纳瓦分校) King Mongkut’s Institute of Technology Ladkrabang(拉差丹awan技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments EMNLP Mains 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03458 2025-10-07 cs.CL 83%

Omni-Embed-Nemotron: A Unified Multimodal Retrieval Model for Text, Image, Audio, and Video

Mengyao Xu, Wenfei Zhou, Yauhen Babakhin, Gabriel Moreira, Ronay Ak, Radek Osmulski, Bo Liu, Even Oldridge, Benedikt Schifferer

机构 * NVIDIA

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25177 2025-09-30 cs.CV 83%

Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding

Bingkui Tong, Jiaer Xia, Kaiyang Zhou

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Hong Kong Baptist University(香港 Baptist大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19994 2025-09-30 cs.CV 83%

Improving Generalizability and Undetectability for Targeted Adversarial Attacks on Multimodal Pre-trained Models

Zhifang Zhang, Jiahan Zhang, Shengjie Zhou, Qi Wei, Shuo He, Feng Liu, Lei Feng

机构 * Southeast University(东南大学) Johns Hopkins University(约翰霍普金斯大学) Chongqing University(重庆大学) Nanyang Technological University(南洋理工大学) University of Melbourne(墨尔本大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18450 2025-09-30 cs.CL 83%

BRIT: Bidirectional Retrieval over Unified Image-Text Graph

Ainulla Khan, Yamada Moyuru, Srinidhi Akella

机构 * Fujitsu Research India(富士通印度研究)

专题命中 跨模态检索 :image-text(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted in EMNLP-2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15472 2025-09-29 cs.CV 83%

Efficient Multimodal Dataset Distillation via Generative Models

Zhenghao Zhao, Haoxuan Wang, Junyi Wu, Yuzhang Shang, Gaowen Liu, Yan Yan

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Central Florida(中央佛罗里达大学) Cisco Research(思科研究)

专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21556 2025-09-29 cs.CL 83%

VAT-KG: Knowledge-Intensive Multimodal Knowledge Graph Dataset for Retrieval-Augmented Generation

Hyeongcheol Park, Jiyoung Seo, MinHyuk Jang, Hogun Park, Ha Dam Baek, Gyusam Chang, Hyeonsoo Im, Sangpil Kim

机构 * Korea University(韩国大学) Sungkyunkwan University(全北大学) Hanwha Systems(韩华系统)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments Project Page: https://vatkg.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16597 2025-09-23 cs.CL 83%

MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models

Luyan Zhang

机构 * Northeastern University(东北大学)

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CL

Comments 13 pages, 6 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17265 2025-09-23 cs.LG cs.AI 83%

SUA: Stealthy Multimodal Large Language Model Unlearning Attack

Xianren Zhang, Hui Liu, Delvin Ce Zhang, Xianfeng Tang, Qi He, Dongwon Lee, Suhang Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Amazon(亚马逊) University of Sheffield(谢菲尔德大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments EMNLP25

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02962 2025-09-18 cs.CV 83%

Resilient Multimodal Industrial Surface Defect Detection with Uncertain Sensors Availability

Shuai Jiang, Yunfeng Ma, Jingyu Zhou, Yuan Bian, Yaonan Wang, Min Liu

机构 * School of Artificial Intelligence and Robotics(人工智能与机器人学院) National Engineering Research Center for Robot Visual Perception and Control Technology(机器人视觉感知与控制技术国家工程研究中心) Hunan University(湖南大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE/ASME Transactions on Mechatronics

Journal ref IEEE/ASME Transactions on Mechatronics, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01275 2025-09-17 cs.AI 83%

Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D

Artemis Panagopoulou, Le Xue, Honglu Zhou, silvio savarese, Ran Xu, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, Juan Carlos Niebles

机构 * Salesforce AI Reseach(Salesforce人工智能研究院) University of Pennsylvania(宾夕法尼亚大学)

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11054 2025-09-16 cs.IT cs.CV math.IT 83%

Rate-Distortion Limits for Multimodal Retrieval: Theory, Optimal Codes, and Finite-Sample Guarantees

Thomas Y. Chen

机构 * Department of Computer Science, Columbia University(计算机科学系,哥伦比亚大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments ICCV MRR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10467 2025-09-16 cs.IR cs.AI cs.CL cs.CV cs.MM 83%

DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph

Mengzheng Yang, Yanfei Ren, David Osei Opoku, Ruochang Li, Peng Ren, Chunxiao Xing

机构 * School of Software, Henan University, Kaifeng 475004, China(河南大学软件学院) BNRist, DCST, RIIT, Tsinghua University, Beijing 100084, China(清华大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 12 pages, 5 figures. Accepted to the 22nd International Conference on Web Information Systems and Applications (WISA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08897 2025-09-12 cs.CV cs.AI cs.CL cs.MM 83%

Recurrence Meets Transformers for Universal Multimodal Retrieval

Davide Caffagni, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

机构 * Department of Education and Humanities, University of Modena and Reggio Emilia(教育与人文学院, Modena and Reggio Emilia大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01360 2025-09-03 cs.CV cs.LG 83%

M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision

Che Liu, Zheng Jiang, Chengyu Fang, Heng Guo, Yan-Jie Zhou, Jiaqi Qu, Le Lu, Minfeng Xu

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Imperial College London(帝国理工学院) Tsinghua University(清华大学) Hupan Lab(华潘实验室)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16701 2025-09-03 cs.IR cs.CL 83%

AlzheimerRAG: Multimodal Retrieval Augmented Generation for Clinical Use Cases using PubMed articles

Aritra Kumar Lahiri, Qinmin Vivian Hu

机构 * Department of Computer Science, Toronto Metropolitan University(计算机科学系,多伦多 Metropolitan 大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Journal ref Machine Learning and Knowledge Extraction. 2025; 7(3):89

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20188 2025-08-29 cs.CV cs.LG 83%

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

Max Torop, Masih Eskandar, Nicholas Kurtansky, Jinyang Liu, Jochen Weber, Octavia Camps, Veronica Rotemberg, Jennifer Dy, Kivanc Kose

机构 * Northeastern University(东北大学) Memorial Sloan Kettering Cancer Center(纪念斯隆凯特琳癌症中心)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20057 2025-08-28 cs.MM 83%

ProMSC-MIS: Prompt-based Multimodal Semantic Communication for Multi-Spectral Image Segmentation

Haoshuo Zhang, Yufei Bo, Meixia Tao

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Comments arXiv admin note: text overlap with arXiv:2508.17920

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02826 2025-08-27 cs.CV 83%

Confidence-driven Gradient Modulation for Multimodal Human Activity Recognition: A Dynamic Contrastive Dual-Path Learning Approach

Panpan Ji, Junni Song, Yifan Lu, Hang Xiao, Hanyu Liu, Chao Li

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13901 2025-08-20 cs.RO cs.CV 83%

Multimodal Data Storage and Retrieval for Embodied AI: A Survey

Yihao Lu, Hao Tang

机构 * School of Economics and Management, South China Normal University(经济管理学院,华南师范大学) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01226 2025-08-05 cs.IR cs.MM 83%

CM$^3$: Calibrating Multimodal Recommendation

Xin Zhou, Yongjie Wang, Zhiqi Shen

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.MM

Comments Working Paper: https://github.com/enoche/CM3

详情

展开后加载摘要…

URL PDF HTML 收藏