arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-23 至 2025-09-23 共收录 103 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 21 篇

2509.00367 2025-09-23 cs.CV 79%

A Multimodal and Multi-centric Head and Neck Cancer Dataset for Segmentation, Diagnosis and Outcome Prediction

Numan Saeed, Salma Hassan, Shahad Hardan, Ahmed Aly, Darya Taratynova, Umair Nawaz, Ufaq Khan, Muhammad Ridzuan, Vincent Andrearczyk, Adrien Depeursinge, Yutong Xie, Thomas Eugene, Raphaël Metz, Mélanie Dore, Gregory Delpon, Vijay Ram Kumar Papineni, Kareem Wahid, Cem Dede, Alaa Mohamed Shawky Ali, Carlos Sjogreen, Mohamed Naser, Clifton D. Fuller, Valentin Oreiller, Mario Jreige, John O. Prior, Catherine Cheze Le Rest, Olena Tankyevych, Pierre Decazes, Su Ruan, Stephanie Tanadini-Lang, Martin Vallières, Hesham Elhalawani, Ronan Abgral, Romain Floch, Kevin Kerleguer, Ulrike Schick, Maelle Mauguen, David Bourhis, Jean-Christophe Leclere, Amandine Sambourg, Arman Rahmim, Mathieu Hatt, Mohammad Yaqub

机构 * Department of Computer Vision, Mohamed bin Zayed University of Artificial Intelligence(计算机视觉系,Mohamed bin Zayed人工智能大学) Department of Machine Learning, Mohamed bin Zayed University of Artificial Intelligence(机器学习系,Mohamed bin Zayed人工智能大学) Institute of Informatics, HES-SO Valais-Wallis University of Applied Sciences and Arts(信息学院,HES-SO瓦莱-杜萨大学应用科学与艺术学院) Department of Nuclear Medicine and Molecular Imaging, Lausanne University Hospital (CHUV)(核医学与分子成像系,洛桑大学医院(CHUV)) Nantes Université, CHU Nantes, Nuclear Medicine Department(南特大学,南特大学医院,核医学部门)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 5 figures. Numan Saeed is the corresponding author. Numan Saeed, Salma Hassan and Shahad Hardan contributed equally to this work. Project page: https://hecktor25.grand-challenge.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04415 2025-09-23 cs.CL 79%

MOMENTS: A Comprehensive Multimodal Benchmark for Theory of Mind

Emilio Villa-Cueva, S M Masrur Ahmed, Rendi Chevi, Jan Christian Blaise Cruz, Kareem Elzeky, Fermin Cristobal, Alham Fikri Aji, Skyler Wang, Rada Mihalcea, Thamar Solorio

机构 * MBZUAI University of Houston(德克萨斯大学休斯顿分校) McGill University(麦吉尔大学) University of Michigan(密歇根大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20168 2025-09-23 cs.CV 79%

Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models

Zhentao He, Can Zhang, Ziheng Wu, Zhenghao Chen, Yufei Zhan, Yifan Li, Zhao Zhang, Xian Wang, Minghui Qiu

机构 * ByteDance(字节跳动) CASIA(中国科学院自动化研究所) RUC(中国人民大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24456 2025-09-23 cs.CL 79%

CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan, Kareem Elzeky, Henok Biadglign Ademtew, Alham Fikri Aji, Vladimir Araujo, Israel Abebe Azime, Jinheon Baek, Frederico Belcavello, Fermin Cristobal, Jan Christian Blaise Cruz, Mary Dabre, Raj Dabre, Toqeer Ehsan, Naome A Etori, Fauzan Farooqui, Jiahui Geng, Guido Ivetta, Thanmay Jayakumar, Soyeong Jeong, Zheng Wei Lim, Aishik Mandal, Sofia Martinelli, Mihail Minkov Mihaylov, Daniil Orel, Aniket Pramanick, Sukannya Purkayastha, Israfel Salazar, Haiyue Song, Tiago Timponi Torrent, Debela Desalegn Yadeta, Injy Hamed, Atnafu Lambebo Tonja, Thamar Solorio

机构 * MBZUAI(人工智能研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16989 2025-09-23 cs.CL 79%

All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark

Davide Testa, Giovanni Bonetta, Raffaella Bernardi, Alessandro Bondielli, Alessandro Lenci, Alessio Miaschi, Lucia Passaro, Bernardo Magnini

机构 * Università di Roma La Sapienza(罗马La Sapienza大学) Fondazione Bruno Kessler (FBK)(布鲁诺·凯斯勒基金会) Free University of Bozen-Bolzano(博兹纳-博尔扎诺自由大学) Dept. of Computer Science, University of Pisa(比萨大学计算机科学系) CoLing Lab, Dept. of Philology, Literature and Linguistics, University of Pisa(比萨大学语言学、文学与语言学系CoLing实验室) Istituto di Linguistica Computazionale "A. Zampolli" (CNR-ILC), ItaliaNLP Lab, Pisa(A. Zampolli计算语言学研究所(CNR-ILC),意大利NLP实验室,比萨)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16517 2025-09-23 cs.CV cs.AI cs.CL cs.MM 77%

Seeing Culture: A Benchmark for Visual Reasoning and Grounding

Burak Satar, Zhixin Ma, Patrick A. Irawan, Wilfried A. Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo

机构 * Singapore Management University(新加坡管理大学) Bandung Institute of Technology(班达理工大学)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Main Conference, https://seeingculture-benchmark.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13542 2025-09-23 cs.MM eess.SP 74%

A multimodal stress detection dataset with facial expressions and physiological signals

Majid Hosseini, Fahad Sohrab, Raju Gottumukkala, Ravi Teja Bhupatiraju, Satya Katragadda, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj

专题命中 多模态评测 :multimodal(title);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20024 2025-09-23 cs.CV cs.AI cs.RO 73%

ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving

Xueyi Liu, Zuodong Zhong, Yuxin Guo, Yun-Fu Liu, Zhiguo Su, Qichao Zhang, Junli Wang, Yinfeng Gao, Yupeng Zheng, Qiao Lin, Huiyong Chen, Dongbin Zhao

机构 * SKL-MAIS, Institute of Automation, Chinese Academy of Sciences, Beijing, China(中国科学院自动化研究所SKL-MAIS部门,北京) School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学人工智能学院,北京) EACON, Fujian, China(福建EACON机构,中国) School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing, China(北京科技大学自动化与电气工程学院)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 18 pages; 9 figures; https://github.com/Liuxueyi/ReasonPlan

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17437 2025-09-23 cs.CL 70%

GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning

Guizhen Chen, Weiwen Xu, Hao Zhang, Hou Pong Chan, Deli Zhao, Anh Tuan Luu, Yu Rong

机构 * Nanyang Technological University(南洋理工大学) DAMO Academy, Alibaba Group(阿里达摩院) Hupan Lab(华普实验室)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CL

Comments Accepted to EMNLP2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19769 2025-09-23 cs.CV cs.LG 70%

BiPrompt-SAM: Enhancing Image Segmentation via Explicit Selection between Point and Text Prompts

Suzhe Xu, Jialin Peng, Chengyuan Zhang

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments metrics went wrong

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17421 2025-09-23 cs.CL cs.MM 62%

RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios

Fei Zhao, Chengqiang Lu, Yufan Shen, Qimeng Wang, Yicheng Qian, Haoxin Zhang, Yan Gao, Yi Wu, Yao Hu, Zhen Wu, Shangyu Xing, Xinyu Dai

机构 * National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Xiaohongshu Inc.(小红书公司) Zhejiang University(浙江大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.MM

Comments Findings of EMNLP 2025 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17418 2025-09-23 cs.CL cs.CV 62%

Vision Language Models Are Not (Yet) Spelling Correctors

Junhong Liang, Bojun Zhang

机构 * MBZUAI Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17353 2025-09-23 cs.AI eess.IV physics.med-ph 57%

Medical AI Consensus: A Multi-Agent Framework for Radiology Report Generation and Evaluation

Ahmed T. Elboardy, Ghada Khoriba, Essam A. Rashed

机构 * Graduate School of Information Science, University of Hyogo(京都大学垣田学园信息科学研究生院) Center for Informatics Science, School of Information Technology and Computer Science, Nile University(尼罗大学信息科学中心)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments NeurIPS2025 Workshop: Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16628 2025-09-23 cs.CV 57%

Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning

Janak Kapuriya, Anwar Shaikh, Arnav Goel, Medha Hira, Apoorv Singh, Jay Saraf, Sanjana, Vaibhav Nauriyal, Avinash Anand, Zhengkui Wang, Rajiv Ratn Shah

机构 * Indraprastha Institute of Information Technology, Delhi(印度德里印度理工学院) Singapore Institute of Technology(新加坡理工学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 5 篇

2507.10571 2025-09-23 cs.AI cs.CL 62%

Agentic AI with Orchestrator-Agent Trust: A Modular Visual Classification Framework with Trust-Aware Orchestration and RAG-Based Reasoning

Konstantinos I. Roumeliotis, Ranjan Sapkota, Manoj Karkee, Nikolaos D. Tselikas

机构 * University of the Peloponnese, Department of Informatics and Telecommunications(希腊皮埃蒙特大学信息与电信系) Cornell University, Department of Biological and Environmental Engineering(康奈尔大学生物与环境工程系)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15293 2025-09-23 cs.CV cs.RO 57%

How Good are Foundation Models in Step-by-Step Embodied Reasoning?

Dinura Dissanayake, Ahmed Heakl, Omkar Thawakar, Noor Ahsan, Ritesh Thawkar, Ketan More, Jean Lahoud, Rao Anwer, Hisham Cholakkal, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan

机构 * Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学) Linköping University(林雪平大学) Australian National University(澳大利亚国立大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments Project page: https://mbzuai-oryx.github.io/FoMER-Bench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17328 2025-09-23 cs.CV cs.HC 57%

UIPro: Unleashing Superior Interaction Capability For GUI Agents

Hongxin Li, Jingran Su, Jingfan Chen, Zheng Ju, Yuntao Chen, Qing Li, Zhaoxiang Zhang

机构 * University of Chinese Academy of Sciences (UCAS)(中国科学院大学) New Laboratory of Pattern Recognition (NLPR), CASIA(中国科学院自动化所模式识别新实验室) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA(中国科学院多模态人工智能系统国家重点实验室) Hong Kong Institute of Science & Innovation, CASIA(中国科学院香港创新科学研究院) PolyU Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17042 2025-09-23 cs.RO 50%

Orchestrate, Generate, Reflect: A VLM-Based Multi-Agent Collaboration Framework for Automated Driving Policy Learning

Zengqi Peng, Yusen Xie, Yubin Wang, Rui Yang, Qifeng Chen, Jun Ma

机构 * Robotics and Autonomous Systems Thrust, The Hong Kong University of Science and Technology (Guangzhou)(机器人与自主系统方向,香港科学与技术大学(广州)) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科学与技术大学) Cheng Kar-Shun Robotics Institute, The Hong Kong University of Science and Technology(陈家骏机器人研究所,香港科学与技术大学)

专题命中 多模态Agent :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16398 2025-09-23 cs.RO cs.LG 50%

Dynamic Objects Relocalization in Changing Environments with Flow Matching

Francesco Argenziano, Miguel Saavedra-Ruiz, Sacha Morin, Daniele Nardi, Liam Paull

机构 * Department of Computer, Automation and Management Engineering, Sapienza University of Rome(罗马大学计算机、自动化与管理工程系) Department of Computer Science and Operations Research, Université de Montréal(蒙特利尔大学计算机科学与运筹学系) Mila - Quebec AI Institute(魁北克人工智能研究所)

专题命中 多模态Agent :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 16 篇

2504.04653 2025-09-23 cs.CV cs.CL 88%

LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts

Yimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof Czarnecki

机构 * University of Waterloo(滑铁卢大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(title);MLLM(abstract);分类 cs.CV、cs.CL

Comments To appear at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17136 2025-09-23 cs.CV cs.AI 84%

SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM

Yuhao Tian, Zheming Yang

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of AI for Industries(工业人工智能研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16900 2025-09-23 cs.CV cs.AI 84%

ME-Mamba: Multi-Expert Mamba with Efficient Knowledge Capture and Fusion for Multimodal Survival Analysis

Chengsheng Zhang, Linhao Qu, Xiaoyu Liu, Zhijian Song

机构 * Digital Medical Research Center, School of Basic Medical Science, Fudan University, Shanghai 200032, China(复旦大学基础医学学院数字医学研究中心) Shanghai Key Lab of Medical Image Computing(上海医学图像计算重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16618 2025-09-23 cs.CV cs.AI 84%

Surgical-MambaLLM: Mamba2-enhanced Multimodal Large Language Model for VQLA in Robotic Surgery

Pengfei Hao, Hongqiu Wang, Shuaibo Li, Zhaohu Xing, Guang Yang, Kaishun Wu, Lei Zhu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Imperial College London(帝国理工学院伦敦分校) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Early accepted by MICCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21059 2025-09-23 cs.CV cs.AI cs.CR cs.LG 81%

FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts

Ziyi Zhang, Zhen Sun, Zongmin Zhang, Jihui Guo, Xinlei He

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17943 2025-09-23 cs.CV cs.LG 79%

Can multimodal representation learning by alignment preserve modality-specific information?

Romain Thoreau, Jessie Levillain, Dawa Derksen

机构 * institutetext(机构文本) CNES(法国国家空间研究中心) INSA-IMT(法国里尔INSA-IMT)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted as a workshop paper at MACLEAN - ECML/PKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17520 2025-09-23 cs.CV 79%

Unified Multimodal Coherent Field: Synchronous Semantic-Spatial-Vision Fusion for Brain Tumor Segmentation

Mingda Zhang, Yuyang Zheng, Ruixiang Tang, Jingru Qiu, Haiyan Ding

机构 * School of Software, Yunnan University(云南大学软件学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17228 2025-09-23 cs.LG cs.CL stat.ME 79%

Causal Representation Learning from Multimodal Clinical Records under Non-Random Modality Missingness

Zihan Liang, Ziwen Pan, Ruoxuan Xiong

机构 * Emory University(埃默里大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments To appear in Proc. of EMNLP 2025 (18 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18005 2025-09-23 cs.RO 78%

M3ET: Efficient Vision-Language Learning for Robotics based on Multimodal Mamba-Enhanced Transformer

Yanxin Zhang, Liang He, Zeyi Kang, Zuheng Ming, Kaixing Zhao

机构 * School of Software Northwestern Polytechnical University Xi'an, China(软件学院 西安理工大学 西安) Laboratoire L2Tl University Sorbonne Paris Nord Paris, France(L2Tl实验室 索邦巴黎北大学 巴黎)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15250 2025-09-23 cs.CV cs.AI 76%

Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token Pruning

Wenda Qin, Andrea Burns, Bryan A. Plummer, Margrit Betke

机构 * Boston University(波士顿大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

Comments Accepted to EMNLP 2025. Data and code to be released at https://github.com/wdqin/VLN-NAP

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17747 2025-09-23 cs.CV cs.AI 73%

Dual-View Alignment Learning with Hierarchical-Prompt for Class-Imbalance Multi-Label Classification

Sheng Huang, Jiexuan Yan, Beiyan Liu, Bo Liu, Richang Hong

机构 * Ministry of Education Key Laboratory of Dependable Service Computing in Cyber Physical Society(教育部可信服务计算网络社会重点实验室) School of Big Data and Software Engineering(大数据与软件工程学院) School of Computer Science and Information Engineering(计算机科学与信息工程学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments accepted by IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏