arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-23 至 2025-09-23 共收录 103 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 11 篇

2502.01960 2025-09-23 cs.LG 88%

MPIC: Position-Independent Multimodal Context Caching System for Efficient MLLM Serving

Shiju Zhao, Junhao Hu, Rongxiao Huang, Jiaqi Zheng, Guihai Chen

机构 * State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) Nanjing University(南京大学) School of Computer Science(计算机学院)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(title,abstract)

Comments 17 pages, 13 figures, the second version

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16597 2025-09-23 cs.CL 83%

MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models

Luyan Zhang

机构 * Northeastern University(东北大学)

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CL

Comments 13 pages, 6 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17265 2025-09-23 cs.LG cs.AI 83%

SUA: Stealthy Multimodal Large Language Model Unlearning Attack

Xianren Zhang, Hui Liu, Delvin Ce Zhang, Xianfeng Tang, Qi He, Dongwon Lee, Suhang Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Amazon(亚马逊) University of Sheffield(谢菲尔德大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments EMNLP25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21105 2025-09-23 cs.IR cs.AI cs.CL 81%

AgentMaster: A Multi-Agent Conversational Framework Using A2A and MCP Protocols for Multimodal Information Retrieval and Analysis

Callie C. Liao, Duoduo Liao, Sai Surya Gadiraju

机构 * Stanford University(斯坦福大学) George Mason University(乔治·马歇尔大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17762 2025-09-23 cs.CV 79%

Neural-MMGS: Multi-modal Neural Gaussian Splats for Large-Scale Scene Reconstruction

Sitian Shen, Georgi Pramatarov, Yifu Tao, Daniele De Martini

机构 * Mobile Robotics Group (MRG), Oxford Robotics Institute, Department of Engineering Science, University of Oxford, UK(移动机器人组(MRG),牛津机器人研究所,工程科学系,牛津大学,英国)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15269 2025-09-23 cs.CV cs.LG 79%

Test-Time Multimodal Backdoor Detection by Contrastive Prompting

Yuwei Niu, Shuo He, Qi Wei, Zongyu Wu, Feng Liu, Lei Feng

机构 * Chongqing University(重庆大学) Nanyang Technological University(南洋理工大学) Penn State University(宾夕法尼亚州立大学) University of Melbourne(墨尔本大学) Southeast University(东南大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to ICML2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14963 2025-09-23 cs.LG 78%

Continual Multimodal Contrastive Learning

Xiaohao Liu, Xiaobo Xia, See-Kiong Ng, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学)

专题命中 跨模态检索 :multimodal(title,abstract)

Comments Accepted by NeurIPS 2025. Codes are available at https://github.com/Xiaohao-Liu/CMCL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13146 2025-09-23 cs.CV cs.LG 70%

Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization

Shuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai, Yueqi Wang, Chan-Wei Hu, Chengxuan Qian, Huaxiu Yao, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯大学) University of Michigan(密歇根大学) UIUC(伊利诺伊大学香槟分校) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Published at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17802 2025-09-23 cs.CV cs.AI 62%

TS-P$^2$CL: Plug-and-Play Dual Contrastive Learning for Vision-Guided Medical Time Series Classification

Qi'ao Xu, Pengfei Wang, Bo Zhong, Tianwen Qian, Xiaoling Wang, Ye Wang, Hong Yu

机构 * East China Normal University(东华大学) Chongqing University of Posts and Telecommunications(重庆邮电大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16157 2025-09-23 cs.CV 57%

Proxy-Embedding as an Adversarial Teacher: An Embedding-Guided Bidirectional Attack for Referring Expression Segmentation Models

Xingbai Chen, Tingchao Fu, Renyang Liu, Wei Zhou, Chao Yi

机构 * National Pilot School of Software, Yunnan University(云南大学软件试点学校) School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院) Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 20pages, 5figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16212 2025-09-23 cs.DB cs.AI 57%

EPIC: Generative AI Platform for Accelerating HPC Operational Data Analytics

Ahmad Maroof Karimi, Woong Shin, Jesse Hines, Tirthankar Ghosal, Naw Safrin Sattar, Feiyi Wang

机构 * National Center for Computational Sciences(国家计算科学中心) Oak Ridge National Laboratory(橡树岭国家实验室)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态生成 12 篇

2509.17589 2025-09-23 cs.AI 83%

Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models

Jun Ling, Yao Qi, Tao Huang, Shibo Zhou, Yanqin Huang, Jiang Yang, Ziqi Song, Ying Zhou, Yang Yang, Heng Tao Shen, Peng Wang

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) Research Center for Scientific Data Hub, Zhejiang Lab, Hangzhou, China(浙江实验室科学数据中心研究中心) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21476 2025-09-23 cs.CV cs.AI 81%

GarmentDiffusion: 3D Garment Sewing Pattern Generation with Multimodal Diffusion Transformers

Xinyu Li, Qi Yao, Yuanda Wang

机构 * Shenfu Research(沈孚研究所) Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments The 34th International Joint Conference on Artificial Intelligence (IJCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16768 2025-09-23 cs.CV 74%

MMPart: Harnessing Multi-Modal Large Language Models for Part-Aware 3D Generation

Omid Bonakdar, Nasser Mozayani

机构 * School of Computer engineering, Iran university of Science and Technology(计算机工程学院,伊朗科学技术大学)

专题命中 多模态生成 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18096 2025-09-23 cs.CV 70%

Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers

Chaehyun Kim, Heeseong Shin, Eunbeen Hong, Heeji Yoon, Anurag Arnab, Paul Hongsuck Seo, Sunghwan Hong, Seungryong Kim

机构 * KAIST AI(韩国科学技术院人工智能研究所) Korea University(韩国大学) ETH Zürich(苏黎世联邦理工学院)

专题命中 多模态生成 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments NeurIPS 2025. Project page: https://cvlab-kaist.github.io/Seg4Diff/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17281 2025-09-23 cs.LG cs.AI cs.CY 57%

Training the next generation of physicians for artificial intelligence-assisted clinical neuroradiology: ASNR MICCAI Brain Tumor Segmentation (BraTS) 2025 Lighthouse Challenge education platform

Raisa Amiruddin, Nikolay Y. Yordanov, Nazanin Maleki, Pascal Fehringer, Athanasios Gkampenis, Anastasia Janas, Kiril Krantchev, Ahmed Moawad, Fabian Umeh, Salma Abosabie, Sara Abosabie, Albara Alotaibi, Mohamed Ghonim, Mohanad Ghonim, Sedra Abou Ali Mhana, Nathan Page, Marko Jakovljevic, Yasaman Sharifi, Prisha Bhatia, Amirreza Manteghinejad, Melisa Guelen, Michael Veronesi, Virginia Hill, Tiffany So, Mark Krycia, Bojan Petrovic, Fatima Memon, Justin Cramer, Elizabeth Schrickel, Vilma Kosovic, Lorenna Vidal, Gerard Thompson, Ichiro Ikuta, Basimah Albalooshy, Ali Nabavizadeh, Nourel Hoda Tahon, Karuna Shekdar, Aashim Bhatia, Claudia Kirsch, Gennaro D'Anna, Philipp Lohmann, Amal Saleh Nour, Andriy Myronenko, Adam Goldman-Yassen, Janet R. Reid, Sanjay Aneja, Spyridon Bakas, Mariam Aboian

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 23 pages, 9 figures, 1 table, 3 supplementary tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21535 2025-09-23 eess.IV cs.CV cs.LG 57%

Exploring the Design Space of 3D MLLMs for CT Report Generation

Mohammed Baharoon, Jun Ma, Congyu Fang, Augustin Toma, Bo Wang

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究所) Department of Biomedical Informatics, Harvard Medical School(生物医学信息学系,哈佛医学院) Peter Munk Cardiac Centre, University Health Network(皮特·蒙克心脏中心,大学健康网络) Medical Biophysics, University of Toronto(医学生物物理系,多伦多大学) Department of Computer Science, University of Toronto(计算机科学系,多伦多大学) Department of Laboratory Medicine and Pathobiology, University of Toronto(实验室医学与病理学系,多伦多大学) AI Hub, University Health Network(人工智能中心,大学健康网络)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16806 2025-09-23 cs.CV 57%

FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation

Fan Yang, Yousong Zhu, Xin Li, Yufei Zhan, Hongyin Zhao, Shurong Zheng, Yaowei Wang, Ming Tang, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Science(中国科学院大学人工智能学院) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室) Wuhan AI Research, Wuhan, China(武汉人工智能研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00455 2025-09-23 cs.RO cs.AI cs.LG 57%

Diffusion Graph Neural Networks and Dataset for Robust Olfactory Navigation in Hazard Robotics

Kordel K. France, Ovidiu Daescu

机构 * Dept. of Computer Science University of Texas at Dallas Richardson, TX, USA(计算机科学系 德州大学达拉斯分校 Richardson, TX, USA)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16436 2025-09-23 cs.CV 57%

Improved mmFormer for Liver Fibrosis Staging via Missing-Modality Compensation

Zhejia Zhang, Junjie Wang, Le Zhang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17850 2025-09-23 cs.RO 50%

SocialTraj: Two-Stage Socially-Aware Trajectory Prediction for Autonomous Driving via Conditional Diffusion Model

Xiao Zhou, Zengqi Peng, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态生成 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17790 2025-09-23 physics.med-ph eess.IV 50%

Conditional Diffusion Models for CT Image Synthesis from CBCT: A Systematic Review

Alzahra Altalib, Chunhui Li, Alessandro Perelli

专题命中 多模态生成 :multi-modal(abstract)

Comments 36 pages, 8 figures, 3 tables, submitted to Elsevier Computerized Medical Imaging and Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17080 2025-09-23 cs.RO 50%

CoPlanner: An Interactive Motion Planner with Contingency-Aware Diffusion for Autonomous Driving

Ruiguo Zhong, Ruoyu Yao, Pei Liu, Xiaolong Chen, Rui Yang, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态生成 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态评测 21 篇

2411.03823 2025-09-23 cs.CV cs.AI cs.CL cs.MM 85%

Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM

Dingjie Song, Sicheng Lai, Mingxuan Wang, Shunian Chen, Lichao Sun, Benyou Wang

机构 * Lehigh University(莱维大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18172 2025-09-23 cs.CL cs.AI 84%

Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering

Zixin Chen, Sicheng Song, Kashun Shum, Yanna Lin, Rui Sheng, Weiqi Wang, Huamin Qu

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments 34 pages in total, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14051 2025-09-23 cs.CV 83%

PROFUSEme: PROstate Cancer Biochemical Recurrence Prediction via FUSEd Multi-modal Embeddings

Suhang You, Carla Pitarch-Abaigar, Sanket Kachole, Sumedh Sonawane, Juhyung Ha, Anish Sudarshan Gada, David Crandall, Rakesh Shiradkar, Spyridon Bakas

机构 * Division of Computational Pathology, Department of Pathology and Laboratory Medicine, Indiana University School of Medicine, Indianapolis, IN, USA(计算病理学部,病理与实验室医学部,印第安纳大学医学院,印第安纳波利斯,印第安纳州,美国) Indiana University Melvin and Bren Simon Comprehensive Cancer Center, Indianapolis, IN, USA(印第安纳大学Melvin和Bren Simon综合癌症中心,印第安纳波利斯,印第安纳州,美国) Luddy School of Informatics, Computing, and Engineering, Indiana University, IN, USA(卢迪信息学、计算与工程学院,印第安纳大学,印第安纳州,美国) Department of Radiology and Imaging Sciences, Indiana University School of Medicine, Indianapolis, IN, USA(放射学与成像科学部,印第安纳大学医学院,印第安纳波利斯,印第安纳州,美国) Department of Biostatistics and Health Data Science, Indiana University School of Medicine, Indianapolis, IN, USA(生物统计学与健康数据科学部,印第安纳大学医学院,印第安纳波利斯,印第安纳州,美国)

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 11 pages, 1 figure, method paper for CHIMERA 2025 Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20685 2025-09-23 cs.LG cs.AI 83%

Progressive Size-Adaptive Federated Learning: A Comprehensive Framework for Heterogeneous Multi-Modal Data Systems

Sajid Hussain, Muhammad Sohail, Nauman Ali Khan, Naima Iltaf, Ihtesham ul Islam

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.AI

Comments Due to some technical issues

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17337 2025-09-23 cs.AI cs.CL 82%

LLaVul: A Multimodal LLM for Interpretable Vulnerability Reasoning about Source Code

Ala Jararweh, Michael Adams, Avinash Sahu, Abdullah Mueen, Afsah Anwar

机构 * Department of Computer Science, The University of New Mexico(计算机科学系,新墨西哥大学) Comprehensive Cancer Center, The University of New Mexico(综合癌症中心,新墨西哥大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Journal ref A. Jararweh, M. Adams, A. Sahu, A. Mueen and A. Anwar, "LLaVul: A Multimodal LLM for Interpretable Vulnerability Reasoning about Source Code," 2025 5th Intelligent Cybersecurity Conference (ICSC), Tampa, FL, USA, 2025, pp. 232-241

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17740 2025-09-23 cs.CV cs.CL 81%

WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification

Yiwen Jiang, Deval Mehta, Siyuan Yan, Yaling Shen, Zimu Wang, Zongyuan Ge

机构 * Faculty of Engineering, Monash University(墨尔本大学工程学院) AIM for Health Lab, Faculty of IT, Monash University(墨尔本大学信息技术学院健康人工智能实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at EMNLP 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17044 2025-09-23 cs.CV 79%

AgriDoctor: A Multimodal Intelligent Assistant for Agriculture

Mingqing Zhang, Zhuoning Xu, Peijie Wang, Rongji Li, Liang Wang, Qiang Liu, Jian Xu, Xuyao Zhang, Shu Wu, Liang Wang

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏