arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9171 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9171 篇

2505.17455 2025-10-21 cs.CL cs.AI 81%

Towards Evaluating Proactive Risk Awareness of Multimodal Language Models

Youliang Yuan, Wenxiang Jiao, Yuejin Xie, Chihao Shen, Menghan Tian, Wenxuan Wang, Jen-tse Huang, Pinjia He

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)数据科学学院) Xiaohongshu Inc.(小红书公司) Renmin University of China(中国人民大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS 2025 (Track on Datasets and Benchmarks)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15684 2025-10-20 cs.CV cs.AI 81%

Towards Label-Free Brain Tumor Segmentation: Unsupervised Learning with Multimodal MRI

Gerard Comas-Quiles, Carles Garcia-Cabrera, Julia Dietlmeier, Noel E. O'Connor, Ferran Marques

机构 * Universitat Politècnica de Catalunya (UPC)(西班牙巴塞罗那理工大学) University College Dublin (UCD)(都柏林大学) Dublin City University (DCU)(都柏林城市大学) Insight Research Ireland Center For Data Analytics(爱尔兰数据分析研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 10 pages, 5 figures, BraTS GoAT 2025 challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08636 2025-10-20 cs.CV cs.AI 81%

Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models

Xingrui Wang, Wufei Ma, Tiezheng Zhang, Celso M de Melo, Jieneng Chen, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Published in CVPR 2025 as Highlight. Data and code are released at https://github.com/XingruiWang/Spatial457

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14958 2025-10-17 cs.CV cs.CL 81%

MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning

Weikang Shi, Aldrich Yu, Rongyao Fang, Houxing Ren, Ke Wang, Aojun Zhou, Changyao Tian, Xinyu Fu, Yuxuan Hu, Zimu Lu, Linjiang Huang, Si Liu, Rui Liu, Hongsheng Li

机构 * Multimedia Laboratory (MMLab), The Chinese University of Hong Kong(中文大学多媒体实验室) Huawei Research(华为研究) BUAA(北京航空学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Project Page: https://mathcanvas.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14307 2025-10-17 cs.CL cs.AI 81%

MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking

Sathyanarayanan Ramamoorthy, Vishwa Shah, Simran Khanuja, Zaid Sheikh, Shan Jie, Ann Chia, Shearman Chua, Graham Neubig

机构 * Carnegie Mellon University(卡内基梅隆大学) Defence Science and Technology Agency(国防科学与技术局)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19330 2025-10-16 eess.SP cs.AI cs.HC cs.LG cs.MM 81%

LibEMER: A novel benchmark and algorithms library for EEG-based Multimodal Emotion Recognition

Zejun Liu, Yunshan Chen, Chengxi Xie, Yugui Xie, Huan Liu

机构 * XJTU-POLIMI Joint School, Xi'an Jiaotong University School of Computer Science Technology, Xi'an Jiaotong University MIGU Video Co., Ltd., Shanghai, China

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI、cs.MM

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10412 2025-10-16 cs.LG cs.AI cs.CL 81%

Time-IMM: A Dataset and Benchmark for Irregular Multimodal Multivariate Time Series

Ching Chang, Jeehyun Hwang, Yidan Shi, Haixin Wang, Wen-Chih Peng, Tien-Fu Chen, Wei Wang

机构 * University of California, Los Angeles(加州大学洛杉矶分校) National Yang Ming Chiao Tung University(国立阳明交通大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments This paper has been accepted by the NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13211 2025-10-16 cs.CV cs.AI 81%

MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance

Subin Kim, Hoonrae Kim, Jihyun Lee, Yejin Jeon, Gary Geunbae Lee

机构 * KT Corporation, Republic of Korea(韩国KT公司) Graduate School of Artificial Intelligence, POSTECH, Republic of Korea(POSTECH人工智能研究生院) Computer Science and Engineering, POSTECH, Republic of Korea(POSTECH计算机科学与工程系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10965 2025-10-14 cs.CL cs.AI 81%

Judge Before Answer: Can MLLM Discern the False Premise in Question?

Jidong Li, Lingyong Fang, Haodong Zhao, Sufeng Duan, Gongshen Liu

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) Inner Mongolia Research Institute, Shanghai Jiao Tong University(上海交通大学内蒙古研究院)

专题命中 多模态评测 :MLLM(title);multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10546 2025-10-14 cs.CV cs.AI 81%

GLOFNet -- A Multimodal Dataset for GLOF Monitoring and Prediction

Zuha Fatima, Muhammad Anser Sohaib, Muhammad Talha, Sidra Sultana, Ayesha Kanwal, Nazia Perwaiz

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08964 2025-10-13 cs.CV cs.CL 81%

Unleashing Perception-Time Scaling to Multimodal Reasoning Models

Yifan Li, Zhenghao Chen, Ziheng Wu, Kun Zhou, Ruipu Luo, Can Zhang, Zhentao He, Yufei Zhan, Wayne Xin Zhao, Minghui Qiu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大模型与智能治理重点实验室) ByteDance(字节跳动) University of California, San Diego(加州大学圣地亚哥分校) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06019 2025-10-10 cs.CV cs.AI eess.IV eess.SP 81%

BRIGHT: A globally distributed multimodal building damage assessment dataset with very-high-resolution for all-weather disaster response

Hongruixuan Chen, Jian Song, Olivier Dietrich, Clifford Broni-Bediako, Weihao Xuan, Junjue Wang, Xinlei Shao, Yimin Wei, Junshi Xia, Cuiling Lan, Konrad Schindler, Naoto Yokoya

机构 * Graduate School of Frontier Sciences, The University of Tokyo(东京大学前沿科学研究生院) RIKEN Center for Advanced Intelligence Project (AIP), RIKEN(日本理化学研究院先进智能项目中心) Department of Photogrammetry and Remote Sensing, ETH Zürich(苏黎世联邦理工学院测绘与遥感系) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03878 2025-10-07 cs.CV cs.AI 81%

Multi-Modal Oral Cancer Detection Using Weighted Ensemble Convolutional Neural Networks

Ajo Babu George, Sreehari J R Ajo Babu George, Sreehari J R Ajo Babu George, Sreehari J R

机构 * Dicemed

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07640 2025-10-07 cs.MM cs.AI 81%

Synthesizing Sentiment-Controlled Feedback For Multimodal Text and Image Data

Puneet Kumar, Sarthak Malik, Balasubramanian Raman, Xiaobai Li

机构 * Center for Machine Vision and Signal Analysis, University of Oulu(机器视觉与信号分析中心,奥卢大学) Indian Institute of Technology Roorkee(印度理工学院罗尔基分校) State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01659 2025-10-03 cs.CL cs.AI 81%

MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization

Yinhong Liu, Jianfeng He, Hang Su, Ruixue Lian, Yi Nian, Jake Vincent, Srikanth Vishnubhotla, Robinson Piramuthu, Saab Mansour

机构 * AWS AI Labs(AWS人工智能实验室) Language Technology Lab, University of Cambridge(语言技术实验室,剑桥大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01904 2025-10-03 cs.CV cs.AI 81%

What are You Looking at? Modality Contribution in Multimodal Medical Deep Learning

Christian Gapp, Elias Tappeiner, Martin Welk, Karl Fritscher, Elke Ruth Gizewski, Rainer Schubert

机构 * Institute of Biomedical Image Analysis UMIT TIROL -- Private University for Health Sciences(生物医学影像分析研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Contribution to Conference for Computer Assisted Radiology and Surgery (CARS 2025)

Journal ref Int J CARS (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24297 2025-10-01 cs.CL cs.AI 81%

Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs

Junying Wang, Zicheng Zhang, Ye Shen, Yalun Wu, Yingji Liang, Yijin Guo, Farong Wen, Wenzhe Li, Xuezhi Zhao, Qi Jia, Guangtao Zhai

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24888 2025-09-30 cs.CV cs.CL 81%

MMRQA: Signal-Enhanced Multimodal Large Language Models for MRI Quality Assessment

Fankai Jia, Daisong Gan, Zhe Zhang, Zhaochi Wen, Chenchen Dan, Dong Liang, Haifeng Wang

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) ShanghaiTech University(上海理工大学) Southern University of Science and Technology(南方科技大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04740 2025-09-30 cs.CV cs.AI 81%

SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data

Samarth Mishra, Kate Saenko, Venkatesh Saligrama

机构 * Boston University(波士顿大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments ICCV 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23267 2025-09-30 cs.CV cs.AI cs.LG 81%

Learning Regional Monsoon Patterns with a Multimodal Attention U-Net

Swaib Ilias Mazumder, Manish Kumar, Aparajita Khan

机构 * 1Computer Science \& Engineering, Indian Institute of Technology Roorkee, India 2Computer Science \& Engineering, Indian Institute of Technology Ropar, India 3 Computer Science \& Engineering, Indian Institute of Technology (BHU) Varanasi, India

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted in Geospatial AI and Applications with Foundation Models (GAIA) 2025, INSAIT and ELLIS Unit Sofia, Bulgaria

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23035 2025-09-30 cs.CV cs.AI 81%

Sensor-Adaptive Flood Mapping with Pre-trained Multi-Modal Transformers across SAR and Multispectral Modalities

Tomohiro Tanaka, Narumasa Tsutsumida

机构 * Graduate school of Science & Engineering, Saitama University, Japan(埼玉大学理工学部)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01565 2025-09-30 cs.CL cs.CV 81%

Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation

Li Zhou, Lutong Yu, Dongchu Xie, Shaohuan Cheng, Wenyan Li, Haizhou Li

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院) Chengdu Technological University(成都理工大学) University of Copenhagen(哥本哈根大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Cultural Analysis, Cultural Visual Understanding, Cultural Image Transcreation. Accepted by EMNLP 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20858 2025-09-26 cs.GR cs.CV cs.MM 81%

ArchGPT: Understanding the World's Architectures with Large Multimodal Models

Yuze Wang, Luo Yang, Junyi Wang, Yue Qi

机构 * State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室) School of Computer Science and Engineering(计算机科学与工程学院) Beihang University(北京航空航天大学) School of Computer Science and Technology(计算机科学与技术学院) Shandong University(山东大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00213 2025-09-26 cs.CV cs.AI 81%

Multimodal Deep Learning for Phyllodes Tumor Classification from Ultrasound and Clinical Data

Farhan Fuad Abir, Abigail Elliott Daly, Kyle Anderman, Tolga Ozmen, Laura J. Brattain

机构 * Department of Electrical and Computer Engineering, University of Central Florida(电子与计算机工程系,中央佛罗里达大学) Massachusetts General Hospital, Department of Surgery, Section of Breast Surgery(麻省总医院,外科部,乳腺外科) Department of Medicine, University of Central Florida College of Medicine(医学系,中央佛罗里达大学医学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments IEEE-EMBS International Conference on Body Sensor Networks (IEEE-EMBS BSN 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19952 2025-09-25 cs.CV cs.AI 81%

When Words Can't Capture It All: Towards Video-Based User Complaint Text Generation with Multimodal Video Complaint Dataset

Sarmistha Das, R E Zera Marveen Lyngkhoi, Kirtan Jain, Vinayak Goyal, Sriparna Saha, Manish Gupta

机构 * Indian Institute of Technology Patna(印度帕纳布理工大学) Microsoft, India(微软印度)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19274 2025-09-24 cs.CL cs.MM 81%

DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture

Arijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka, Nemil Shah, Abhilekh Borah, Vanshika Shah, Nishant Mishra, Sriparna Saha

机构 * Indian Institute of Technology Patna(印度理工学院帕纳巴分校) Banasthali Vidyapeeth University(班纳萨利大学) Pandit Deendayal Energy University(德英德能源大学) Manipal University Jaipur(马哈拉施特拉邦大学贾伊普尔分校) Dwarkadas J. Sanghvi College of Engineering(德瓦尔卡斯J.桑格维工程学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments EMNLP MAINS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17740 2025-09-23 cs.CV cs.CL 81%

WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification

Yiwen Jiang, Deval Mehta, Siyuan Yan, Yaling Shen, Zimu Wang, Zongyuan Ge

机构 * Faculty of Engineering, Monash University(墨尔本大学工程学院) AIM for Health Lab, Faculty of IT, Monash University(墨尔本大学信息技术学院健康人工智能实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at EMNLP 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15701 2025-09-22 cs.CL cs.SD eess.AS 81%

Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment

Ke Wang, Wenning Wei, Yan Deng, Lei He, Sheng Zhao

机构 * Microsoft, Beijing, China(微软北京研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、eess.AS

Comments submitted to ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10059 2025-09-15 cs.CV cs.AI 81%

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration

Yue Zhou, Litong Feng, Mengcheng Lan, Xue Yang, Qingyun Li, Yiping Ke, Xue Jiang, Wayne Zhang

机构 * Department of Physics, J.K. Institute of Science(J.K.科学研究院物理系) World Scientific University(世界科学大学) University of Intelligent Studies(智能研究大学) East China Normal University(华东师范大学) Nanyang Technological University(南洋理工大学) SenseTime Research(商汤科技研究院) Shanghai Jiao Tong University(上海交通大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 17 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09254 2025-09-12 cs.CV cs.MM 81%

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis

Jing Hao, Yuxuan Fan, Yanpeng Sun, Kaixin Guo, Lizhuo Lin, Jinrong Yang, Qi Yong H. Ai, Lun M. Wong, Hao Tang, Kuo Feng Hung

机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院) The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学) National University of Singapore(新加坡国立大学) CVTE Sun Yat-sen University(孙中山大学) Department of Diagnostic Radiology, The University of Hong Kong(香港大学放射科) Imaging and Interventional Radiology, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院影像与介入放射科) School of Computer Science, Peking University(北京大学计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 40 pages, 26 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏