arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9176 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9176 篇

2509.13857 2025-09-30 cs.RO cs.CV 79%

InterKey: Cross-modal Intersection Keypoints for Global Localization on OpenStreetMap

Nguyen Hoang Khoi Tran, Julie Stephany Berrio, Mao Shan, Stewart Worrall

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23811 2025-09-30 cs.AI 79%

AnveshanaAI: A Multimodal Platform for Adaptive AI/ML Education through Automated Question Generation and Interactive Assessment

Rakesh Thakur, Diksha Khandelwal, Shreya Tiwari

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 11 pages, 12 figures. Under review as a conference paper at ICLR 2026. Preprint version posted on arXiv

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23783 2025-09-30 cs.AI 79%

Falcon: A Cross-Modal Evaluation Dataset for Comprehensive Safety Perception

Qi Xue, Minrui Jiang, Runjia Zhang, Xiurui Xie, Pei Ke, Guisong Liu

机构 * Laboratory of Intelligent Collaborative Computing University of Electronic Science and Technology of China(智能协同计算实验室 电子科技大学)

专题命中 多模态评测 :cross-modal(title);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23715 2025-09-30 cs.CL cs.LG 79%

Do LLMs Understand Romanian Driving Laws? A Study on Multimodal and Fine-Tuned Question Answering

Eduard Barbu, Adrian Marius Dumitran

机构 * Faculty of Mathematics and Informatics, University of Bucharest(数学与信息学系,布加勒斯特大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted@ CONSILR 2025 Bucharest Romania 9-10 October

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22810 2025-09-30 eess.SP cs.CV 79%

Introducing Multimodal Paradigm for Learning Sleep Staging PSG via General-Purpose Model

Jianheng Zhou, Chenyu Liu, Jinan Zhou, Yi Ding, Yang Liu, Haoran Luo, Ziyu Jia, Xinliang Zhou

机构 * School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(南洋理工大学电子与电气工程学院) College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算与数据科学学院) Nutanix, CA, USA(Nutanix公司) Center for Machine Vision and Signal Analysis, University of Oulu, Oulu, Finland(奥卢大学机器视觉与信号分析中心) Institute of Automation, Chinese Academy of Sciences, Beijing, China(中国科学院自动化研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21143 2025-09-30 cs.RO cs.CL 79%

Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems

Junfeng Yan, Biao Wu, Meng Fang, Ling Chen

机构 * Australian Artificial Intelligence Institute(澳大利亚人工智能研究所) University of Liverpool(利物浦大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments 10 pages, 5 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21979 2025-09-30 cs.CL 79%

Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset

Fakhraddin Alwajih, Samar M. Magdy, Abdellah El Mekki, Omer Nacar, Youssef Nafea, Safaa Taher Abdelfadil, Abdulfattah Mohammed Yahya, Hamzah Luqman, Nada Almarwani, Samah Aloufi, Baraah Qawasmen, Houdaifa Atou, Serry Sibaee, Hamzah A. Alsayadi, Walid Al-Dhabyani, Maged S. Al-shaibani, Aya El Aatar, Nour Qandos, Rahaf Alhamouri, Samar Ahmad, Mohammed Anwar Al-Ghrawi, Aminetou Yacoub, Ruwa AbuHweidi, Vatimetou Mohamed Lemin, Reem Abdel-Salam, Ahlam Bashiti, Aisha Alansari, Ahmed Ashraf, Nora Alturayeif, Alcides Alcoba Inciarte, Adel Ammar, Abdelrahim A. Elmadany, Mohamedou Cheikh Tourad, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments https://github.com/UBC-NLP/pearl

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22377 2025-09-29 cs.CV 79%

Effectiveness of Large Multimodal Models in Detecting Disinformation: Experimental Results

Yasmina Kheddache, Marc Lalonde

机构 * Département d’informatique et recherche opérationnelle (D.I.R.O.) Université de Montréal Pavillon André-Aisenstadt 2920, chemin de la Tour Montréal (QC) H3T 1N8(蒙特利尔大学计算机与运筹学系) R&D Dept. Computer Research Institute of Montreal (CRIM) 405 Ogilvy Ave., #101 Montreal, Qc, Canada(蒙特利尔计算机研究 institute)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22019 2025-09-29 cs.CV 79%

EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking

Yuki Sakai, Ryosuke Furuta, Juichun Yen, Yoichi Sato

机构 * The University of Tokyo(东京大学)

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments Accepted to the I-HFM Workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21805 2025-09-29 cs.CL 79%

Towards Minimal Causal Representations for Human Multimodal Language Understanding

Menghua Jiang, Yuncheng Jiang, Haifeng Hu, Sijie Mai

机构 * School of Computer Science, South China Normal University(华南师范大学计算机学院) School of Electronics and Information Technology, Sun Yat-sen University(中山大学电子与信息学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21600 2025-09-29 cs.AI cs.LG 79%

Automated and Interpretable Survival Analysis from Multimodal Data

Mafalda Malafaia, Peter A. N. Bosman, Coen Rasch, Tanja Alderliesten

机构 * Centrum Wiskunde & Informatica(数学与信息学研究中心) Delft University of Technology(代尔夫特理工大学) Leiden University Medical Center(莱顿大学医学中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 4 figures; 4 tables; 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16044 2025-09-29 eess.AS cs.LG eess.IV eess.SP 79%

Multimodal Biomarkers for Schizophrenia: Towards Individual Symptom Severity Estimation

Gowtham Premananth, Philip Resnik, Sonia Bansal, Deanna L. Kelly, Carol Espy-Wilson

机构 * Department of Electrical \& Computer Engineering Institute for Advanced Computer Studies School of Medicine

专题命中 多模态评测 :multimodal(title,abstract);分类 eess.AS

Comments Accepted to be presented at Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21352 2025-09-29 cs.CV cs.LG 79%

Improving Autism Detection with Multimodal Behavioral Analysis

William Saakyan, Matthias Norden, Lola Eversmann, Simon Kirsch, Muyu Lin, Simon Guendelman, Isabel Dziobek, Hanna Drimalla

机构 * Center for Cognitive Interaction Technology (CITEC), Bielefeld University(认知交互技术中心(CITEC),比勒菲尔德大学) Institute of Psychology, Humboldt University of Berlin(心理学研究所,洪堡大学) Department of Psychiatry and Psychotherapy, Medical Center-University of Freiburg(精神病学与心理治疗系,弗赖堡医学院-大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19096 2025-09-26 cs.CV cs.SE 79%

Investigating Traffic Accident Detection Using Multimodal Large Language Models

Ilhan Skender, Kailin Tong, Selim Solmaz, Daniel Watzenig

机构 * Embedded Systems Group (Dept.-E)(嵌入式系统组) Virtual Vehicle Research GmbH(虚拟车辆研究公司) Control Systems Group (Dept.-E)(控制系统组) Institute of Visual Computing(视觉计算研究所) Graz University of Technology(格拉茨技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted for presentation at the 2025 IEEE International Automated Vehicle Validation Conference (IAVVC 2025). Final version to appear in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23601 2025-09-25 cs.CV 79%

EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis

Shengyuan Liu, Boyun Zheng, Wenting Chen, Zhihao Peng, Zhenfei Yin, Jing Shao, Jiancong Hu, Yixuan Yuan

机构 * Chinese University of Hong Kong(中国香港大学) City University of Hong Kong(香港城市大学) University of Oxford(牛津大学) Shanghai AI Laboratory(上海人工智能实验室) The Sixth Affiliated Hospital, Sun Yat-sen University(中山大学第六附属医院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments 40 pages, 22 figures; Accepted by NeurIPS 2025 Dataset and Benchmark Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02876 2025-09-25 cs.CV cs.LG 79%

Multimodal Reference Visual Grounding

Yangxiao Lu, Ruosen Li, Liqiang Jing, Jikai Wang, Xinya Du, Yunhui Guo, Nicholas Ruozzi, Yu Xiang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Project page with our code and dataset: https://irvlutd.github.io/MultiGrounding

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17044 2025-09-23 cs.CV 79%

AgriDoctor: A Multimodal Intelligent Assistant for Agriculture

Mingqing Zhang, Zhuoning Xu, Peijie Wang, Rongji Li, Liang Wang, Qiang Liu, Jian Xu, Xuyao Zhang, Shu Wu, Liang Wang

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00367 2025-09-23 cs.CV 79%

A Multimodal and Multi-centric Head and Neck Cancer Dataset for Segmentation, Diagnosis and Outcome Prediction

Numan Saeed, Salma Hassan, Shahad Hardan, Ahmed Aly, Darya Taratynova, Umair Nawaz, Ufaq Khan, Muhammad Ridzuan, Vincent Andrearczyk, Adrien Depeursinge, Yutong Xie, Thomas Eugene, Raphaël Metz, Mélanie Dore, Gregory Delpon, Vijay Ram Kumar Papineni, Kareem Wahid, Cem Dede, Alaa Mohamed Shawky Ali, Carlos Sjogreen, Mohamed Naser, Clifton D. Fuller, Valentin Oreiller, Mario Jreige, John O. Prior, Catherine Cheze Le Rest, Olena Tankyevych, Pierre Decazes, Su Ruan, Stephanie Tanadini-Lang, Martin Vallières, Hesham Elhalawani, Ronan Abgral, Romain Floch, Kevin Kerleguer, Ulrike Schick, Maelle Mauguen, David Bourhis, Jean-Christophe Leclere, Amandine Sambourg, Arman Rahmim, Mathieu Hatt, Mohammad Yaqub

机构 * Department of Computer Vision, Mohamed bin Zayed University of Artificial Intelligence(计算机视觉系,Mohamed bin Zayed人工智能大学) Department of Machine Learning, Mohamed bin Zayed University of Artificial Intelligence(机器学习系,Mohamed bin Zayed人工智能大学) Institute of Informatics, HES-SO Valais-Wallis University of Applied Sciences and Arts(信息学院,HES-SO瓦莱-杜萨大学应用科学与艺术学院) Department of Nuclear Medicine and Molecular Imaging, Lausanne University Hospital (CHUV)(核医学与分子成像系,洛桑大学医院(CHUV)) Nantes Université, CHU Nantes, Nuclear Medicine Department(南特大学,南特大学医院,核医学部门)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 5 figures. Numan Saeed is the corresponding author. Numan Saeed, Salma Hassan and Shahad Hardan contributed equally to this work. Project page: https://hecktor25.grand-challenge.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04415 2025-09-23 cs.CL 79%

MOMENTS: A Comprehensive Multimodal Benchmark for Theory of Mind

Emilio Villa-Cueva, S M Masrur Ahmed, Rendi Chevi, Jan Christian Blaise Cruz, Kareem Elzeky, Fermin Cristobal, Alham Fikri Aji, Skyler Wang, Rada Mihalcea, Thamar Solorio

机构 * MBZUAI University of Houston(德克萨斯大学休斯顿分校) McGill University(麦吉尔大学) University of Michigan(密歇根大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20168 2025-09-23 cs.CV 79%

Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models

Zhentao He, Can Zhang, Ziheng Wu, Zhenghao Chen, Yufei Zhan, Yifan Li, Zhao Zhang, Xian Wang, Minghui Qiu

机构 * ByteDance(字节跳动) CASIA(中国科学院自动化研究所) RUC(中国人民大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24456 2025-09-23 cs.CL 79%

CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan, Kareem Elzeky, Henok Biadglign Ademtew, Alham Fikri Aji, Vladimir Araujo, Israel Abebe Azime, Jinheon Baek, Frederico Belcavello, Fermin Cristobal, Jan Christian Blaise Cruz, Mary Dabre, Raj Dabre, Toqeer Ehsan, Naome A Etori, Fauzan Farooqui, Jiahui Geng, Guido Ivetta, Thanmay Jayakumar, Soyeong Jeong, Zheng Wei Lim, Aishik Mandal, Sofia Martinelli, Mihail Minkov Mihaylov, Daniil Orel, Aniket Pramanick, Sukannya Purkayastha, Israfel Salazar, Haiyue Song, Tiago Timponi Torrent, Debela Desalegn Yadeta, Injy Hamed, Atnafu Lambebo Tonja, Thamar Solorio

机构 * MBZUAI(人工智能研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16989 2025-09-23 cs.CL 79%

All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark

Davide Testa, Giovanni Bonetta, Raffaella Bernardi, Alessandro Bondielli, Alessandro Lenci, Alessio Miaschi, Lucia Passaro, Bernardo Magnini

机构 * Università di Roma La Sapienza(罗马La Sapienza大学) Fondazione Bruno Kessler (FBK)(布鲁诺·凯斯勒基金会) Free University of Bozen-Bolzano(博兹纳-博尔扎诺自由大学) Dept. of Computer Science, University of Pisa(比萨大学计算机科学系) CoLing Lab, Dept. of Philology, Literature and Linguistics, University of Pisa(比萨大学语言学、文学与语言学系CoLing实验室) Istituto di Linguistica Computazionale "A. Zampolli" (CNR-ILC), ItaliaNLP Lab, Pisa(A. Zampolli计算语言学研究所(CNR-ILC),意大利NLP实验室,比萨)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15839 2025-09-22 cs.CL 79%

Multi-Physics: A Comprehensive Benchmark for Multimodal LLMs Reasoning on Chinese Multi-Subject Physics Problems

Zhongze Luo, Zhenshuai Yin, Yongxin Guo, Zhichao Wang, Jionghao Zhu, Xiaoying Tang

机构 * School of Science(科学学院) Engineering, The Chinese University of Hong Kong, Shenzhen, China(工程学院,香港中文大学(深圳))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13212 2025-09-22 cs.CV 79%

Semantic Change Detection of Roads and Bridges: A Fine-grained Dataset and Multimodal Frequency-driven Detector

Qingling Shu, Sibao Chen, Xiao Wang, Zhihui You, Wei Lu, Jin Tang, Bin Luo

机构 * MOE Key Lab of ICSP, IMIS Lab of Anhui Province, Anhui Provincial Key Lab of Multimodal Cognitive Computation, School of Computer Science and Technology, Anhui University, Hefei, China(教育部ICSP重点实验室、安徽IMIS实验室、安徽省多模态认知计算重点实验室、计算机科学与技术学院、安徽大学,合肥,中国) School of Public Safety and Emergency Management, Anhui University of Science and Technology, Hefei, China(公共安全与应急管理学院、安徽理工大学,合肥,中国)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13289 2025-09-17 cs.CV eess.IV 79%

Image Realness Assessment and Localization with Multimodal Features

Lovish Kaushik, Agnij Biswas, Somdyuti Paul

机构 * Indian Institute of Technology, Kharagpur(印度理工学院,克哈拉格普)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12963 2025-09-17 cs.CV cs.LG 79%

MMMS: Multi-Modal Multi-Surface Interactive Segmentation

Robin Schön, Julian Lorenz, Katja Ludwig, Daniel Kienzle, Rainer Lienhart

机构 * Fakultät für Angewandte Informatik University of Augsburg(应用信息学院乌尔姆大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments 19 pages, 11 figures, 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12750 2025-09-17 cs.CV 79%

What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference Alignment

Rishab Parthasarathy, Jasmine Collins, Cory Stephenson

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 7 pages, 9 figures, 3 tables; appendix 16 pages, 9 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12287 2025-09-17 eess.IV cs.CV cs.LG 79%

Enhancing Radiographic Disease Detection with MetaCheX, a Context-Aware Multimodal Model

Nathan He, Cody Chen

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments All authors contributed equally, 5 pages, 2 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19075 2025-09-17 cs.CV 79%

HoloDx: Knowledge- and Data-Driven Multimodal Diagnosis of Alzheimer's Disease

Qiuhui Chen, Jintao Wang, Gang Wang, Yi Hong

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) Department of Neurology, Renji Hospital Affiliated to Shanghai Jiao Tong University School of Medicine(上海交通大学医学院附属仁济医院神经内科)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Medical Imaging (TMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11620 2025-09-16 cs.CL cs.CY 79%

AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment

Kun Li, Lai-Man Po, Hongzheng Yang, Xuyuan Xu, Kangcheng Liu, Yuzhi Zhao

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏