arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-23 至 2025-09-23 共收录 21 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 21 篇

2411.03823 2025-09-23 cs.CV cs.AI cs.CL cs.MM 85%

Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM

Dingjie Song, Sicheng Lai, Mingxuan Wang, Shunian Chen, Lichao Sun, Benyou Wang

机构 * Lehigh University(莱维大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18172 2025-09-23 cs.CL cs.AI 84%

Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering

Zixin Chen, Sicheng Song, Kashun Shum, Yanna Lin, Rui Sheng, Weiqi Wang, Huamin Qu

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments 34 pages in total, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14051 2025-09-23 cs.CV 83%

PROFUSEme: PROstate Cancer Biochemical Recurrence Prediction via FUSEd Multi-modal Embeddings

Suhang You, Carla Pitarch-Abaigar, Sanket Kachole, Sumedh Sonawane, Juhyung Ha, Anish Sudarshan Gada, David Crandall, Rakesh Shiradkar, Spyridon Bakas

机构 * Division of Computational Pathology, Department of Pathology and Laboratory Medicine, Indiana University School of Medicine, Indianapolis, IN, USA(计算病理学部,病理与实验室医学部,印第安纳大学医学院,印第安纳波利斯,印第安纳州,美国) Indiana University Melvin and Bren Simon Comprehensive Cancer Center, Indianapolis, IN, USA(印第安纳大学Melvin和Bren Simon综合癌症中心,印第安纳波利斯,印第安纳州,美国) Luddy School of Informatics, Computing, and Engineering, Indiana University, IN, USA(卢迪信息学、计算与工程学院,印第安纳大学,印第安纳州,美国) Department of Radiology and Imaging Sciences, Indiana University School of Medicine, Indianapolis, IN, USA(放射学与成像科学部,印第安纳大学医学院,印第安纳波利斯,印第安纳州,美国) Department of Biostatistics and Health Data Science, Indiana University School of Medicine, Indianapolis, IN, USA(生物统计学与健康数据科学部,印第安纳大学医学院,印第安纳波利斯,印第安纳州,美国)

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 11 pages, 1 figure, method paper for CHIMERA 2025 Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20685 2025-09-23 cs.LG cs.AI 83%

Progressive Size-Adaptive Federated Learning: A Comprehensive Framework for Heterogeneous Multi-Modal Data Systems

Sajid Hussain, Muhammad Sohail, Nauman Ali Khan, Naima Iltaf, Ihtesham ul Islam

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.AI

Comments Due to some technical issues

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17337 2025-09-23 cs.AI cs.CL 82%

LLaVul: A Multimodal LLM for Interpretable Vulnerability Reasoning about Source Code

Ala Jararweh, Michael Adams, Avinash Sahu, Abdullah Mueen, Afsah Anwar

机构 * Department of Computer Science, The University of New Mexico(计算机科学系,新墨西哥大学) Comprehensive Cancer Center, The University of New Mexico(综合癌症中心,新墨西哥大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Journal ref A. Jararweh, M. Adams, A. Sahu, A. Mueen and A. Anwar, "LLaVul: A Multimodal LLM for Interpretable Vulnerability Reasoning about Source Code," 2025 5th Intelligent Cybersecurity Conference (ICSC), Tampa, FL, USA, 2025, pp. 232-241

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17740 2025-09-23 cs.CV cs.CL 81%

WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification

Yiwen Jiang, Deval Mehta, Siyuan Yan, Yaling Shen, Zimu Wang, Zongyuan Ge

机构 * Faculty of Engineering, Monash University(墨尔本大学工程学院) AIM for Health Lab, Faculty of IT, Monash University(墨尔本大学信息技术学院健康人工智能实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at EMNLP 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17044 2025-09-23 cs.CV 79%

AgriDoctor: A Multimodal Intelligent Assistant for Agriculture

Mingqing Zhang, Zhuoning Xu, Peijie Wang, Rongji Li, Liang Wang, Qiang Liu, Jian Xu, Xuyao Zhang, Shu Wu, Liang Wang

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00367 2025-09-23 cs.CV 79%

A Multimodal and Multi-centric Head and Neck Cancer Dataset for Segmentation, Diagnosis and Outcome Prediction

Numan Saeed, Salma Hassan, Shahad Hardan, Ahmed Aly, Darya Taratynova, Umair Nawaz, Ufaq Khan, Muhammad Ridzuan, Vincent Andrearczyk, Adrien Depeursinge, Yutong Xie, Thomas Eugene, Raphaël Metz, Mélanie Dore, Gregory Delpon, Vijay Ram Kumar Papineni, Kareem Wahid, Cem Dede, Alaa Mohamed Shawky Ali, Carlos Sjogreen, Mohamed Naser, Clifton D. Fuller, Valentin Oreiller, Mario Jreige, John O. Prior, Catherine Cheze Le Rest, Olena Tankyevych, Pierre Decazes, Su Ruan, Stephanie Tanadini-Lang, Martin Vallières, Hesham Elhalawani, Ronan Abgral, Romain Floch, Kevin Kerleguer, Ulrike Schick, Maelle Mauguen, David Bourhis, Jean-Christophe Leclere, Amandine Sambourg, Arman Rahmim, Mathieu Hatt, Mohammad Yaqub

机构 * Department of Computer Vision, Mohamed bin Zayed University of Artificial Intelligence(计算机视觉系,Mohamed bin Zayed人工智能大学) Department of Machine Learning, Mohamed bin Zayed University of Artificial Intelligence(机器学习系,Mohamed bin Zayed人工智能大学) Institute of Informatics, HES-SO Valais-Wallis University of Applied Sciences and Arts(信息学院,HES-SO瓦莱-杜萨大学应用科学与艺术学院) Department of Nuclear Medicine and Molecular Imaging, Lausanne University Hospital (CHUV)(核医学与分子成像系,洛桑大学医院(CHUV)) Nantes Université, CHU Nantes, Nuclear Medicine Department(南特大学,南特大学医院,核医学部门)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 5 figures. Numan Saeed is the corresponding author. Numan Saeed, Salma Hassan and Shahad Hardan contributed equally to this work. Project page: https://hecktor25.grand-challenge.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04415 2025-09-23 cs.CL 79%

MOMENTS: A Comprehensive Multimodal Benchmark for Theory of Mind

Emilio Villa-Cueva, S M Masrur Ahmed, Rendi Chevi, Jan Christian Blaise Cruz, Kareem Elzeky, Fermin Cristobal, Alham Fikri Aji, Skyler Wang, Rada Mihalcea, Thamar Solorio

机构 * MBZUAI University of Houston(德克萨斯大学休斯顿分校) McGill University(麦吉尔大学) University of Michigan(密歇根大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20168 2025-09-23 cs.CV 79%

Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models

Zhentao He, Can Zhang, Ziheng Wu, Zhenghao Chen, Yufei Zhan, Yifan Li, Zhao Zhang, Xian Wang, Minghui Qiu

机构 * ByteDance(字节跳动) CASIA(中国科学院自动化研究所) RUC(中国人民大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24456 2025-09-23 cs.CL 79%

CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan, Kareem Elzeky, Henok Biadglign Ademtew, Alham Fikri Aji, Vladimir Araujo, Israel Abebe Azime, Jinheon Baek, Frederico Belcavello, Fermin Cristobal, Jan Christian Blaise Cruz, Mary Dabre, Raj Dabre, Toqeer Ehsan, Naome A Etori, Fauzan Farooqui, Jiahui Geng, Guido Ivetta, Thanmay Jayakumar, Soyeong Jeong, Zheng Wei Lim, Aishik Mandal, Sofia Martinelli, Mihail Minkov Mihaylov, Daniil Orel, Aniket Pramanick, Sukannya Purkayastha, Israfel Salazar, Haiyue Song, Tiago Timponi Torrent, Debela Desalegn Yadeta, Injy Hamed, Atnafu Lambebo Tonja, Thamar Solorio

机构 * MBZUAI(人工智能研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16989 2025-09-23 cs.CL 79%

All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark

Davide Testa, Giovanni Bonetta, Raffaella Bernardi, Alessandro Bondielli, Alessandro Lenci, Alessio Miaschi, Lucia Passaro, Bernardo Magnini

机构 * Università di Roma La Sapienza(罗马La Sapienza大学) Fondazione Bruno Kessler (FBK)(布鲁诺·凯斯勒基金会) Free University of Bozen-Bolzano(博兹纳-博尔扎诺自由大学) Dept. of Computer Science, University of Pisa(比萨大学计算机科学系) CoLing Lab, Dept. of Philology, Literature and Linguistics, University of Pisa(比萨大学语言学、文学与语言学系CoLing实验室) Istituto di Linguistica Computazionale "A. Zampolli" (CNR-ILC), ItaliaNLP Lab, Pisa(A. Zampolli计算语言学研究所(CNR-ILC),意大利NLP实验室,比萨)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16517 2025-09-23 cs.CV cs.AI cs.CL cs.MM 77%

Seeing Culture: A Benchmark for Visual Reasoning and Grounding

Burak Satar, Zhixin Ma, Patrick A. Irawan, Wilfried A. Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo

机构 * Singapore Management University(新加坡管理大学) Bandung Institute of Technology(班达理工大学)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Main Conference, https://seeingculture-benchmark.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13542 2025-09-23 cs.MM eess.SP 74%

A multimodal stress detection dataset with facial expressions and physiological signals

Majid Hosseini, Fahad Sohrab, Raju Gottumukkala, Ravi Teja Bhupatiraju, Satya Katragadda, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj

专题命中 多模态评测 :multimodal(title);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20024 2025-09-23 cs.CV cs.AI cs.RO 73%

ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving

Xueyi Liu, Zuodong Zhong, Yuxin Guo, Yun-Fu Liu, Zhiguo Su, Qichao Zhang, Junli Wang, Yinfeng Gao, Yupeng Zheng, Qiao Lin, Huiyong Chen, Dongbin Zhao

机构 * SKL-MAIS, Institute of Automation, Chinese Academy of Sciences, Beijing, China(中国科学院自动化研究所SKL-MAIS部门,北京) School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学人工智能学院,北京) EACON, Fujian, China(福建EACON机构,中国) School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing, China(北京科技大学自动化与电气工程学院)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 18 pages; 9 figures; https://github.com/Liuxueyi/ReasonPlan

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17437 2025-09-23 cs.CL 70%

GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning

Guizhen Chen, Weiwen Xu, Hao Zhang, Hou Pong Chan, Deli Zhao, Anh Tuan Luu, Yu Rong

机构 * Nanyang Technological University(南洋理工大学) DAMO Academy, Alibaba Group(阿里达摩院) Hupan Lab(华普实验室)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CL

Comments Accepted to EMNLP2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19769 2025-09-23 cs.CV cs.LG 70%

BiPrompt-SAM: Enhancing Image Segmentation via Explicit Selection between Point and Text Prompts

Suzhe Xu, Jialin Peng, Chengyuan Zhang

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments metrics went wrong

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17421 2025-09-23 cs.CL cs.MM 62%

RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios

Fei Zhao, Chengqiang Lu, Yufan Shen, Qimeng Wang, Yicheng Qian, Haoxin Zhang, Yan Gao, Yi Wu, Yao Hu, Zhen Wu, Shangyu Xing, Xinyu Dai

机构 * National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Xiaohongshu Inc.(小红书公司) Zhejiang University(浙江大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.MM

Comments Findings of EMNLP 2025 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17418 2025-09-23 cs.CL cs.CV 62%

Vision Language Models Are Not (Yet) Spelling Correctors

Junhong Liang, Bojun Zhang

机构 * MBZUAI Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17353 2025-09-23 cs.AI eess.IV physics.med-ph 57%

Medical AI Consensus: A Multi-Agent Framework for Radiology Report Generation and Evaluation

Ahmed T. Elboardy, Ghada Khoriba, Essam A. Rashed

机构 * Graduate School of Information Science, University of Hyogo(京都大学垣田学园信息科学研究生院) Center for Informatics Science, School of Information Technology and Computer Science, Nile University(尼罗大学信息科学中心)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments NeurIPS2025 Workshop: Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16628 2025-09-23 cs.CV 57%

Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning

Janak Kapuriya, Anwar Shaikh, Arnav Goel, Medha Hira, Apoorv Singh, Jay Saraf, Sanjana, Vaibhav Nauriyal, Avinash Anand, Zhengkui Wang, Rajiv Ratn Shah

机构 * Indraprastha Institute of Information Technology, Delhi(印度德里印度理工学院) Singapore Institute of Technology(新加坡理工学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏