arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9176 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9176 篇

2412.02734 2025-07-16 cs.CV cs.RO 79%

MVCTrack: Boosting 3D Point Cloud Tracking via Multimodal-Guided Virtual Cues

Zhaofeng Hu, Sifan Zhou, Zhihang Yuan, Dawei Yang, Shibo Zhao, Ci-Jyun Liang

机构 * Stony Brook University(石英溪大学) Southeast University(东南大学) Houmo AI Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09009 2025-07-15 cs.LG cs.AI 79%

Multimodal Cardiovascular Risk Profiling Using Self-Supervised Learning of Polysomnography

Zhengxiao He, Huayu Li, Geng Yuan, William D. S. Killgore, Stuart F. Quan, Chen X. Chen, Ao Li

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10563 2025-07-15 cs.CV 79%

MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks

Jiacheng Chen, Tianhao Liang, Sherman Siu, Zhengqing Wang, Kai Wang, Yubo Wang, Yuansheng Ni, Wang Zhu, Ziyan Jiang, Bohan Lyu, Dongfu Jiang, Xuan He, Yuan Liu, Hexiang Hu, Xiang Yue, Wenhu Chen

机构 * Core Contributors(核心贡献者) Tiger-AI-Lab(虎鲸人工智能实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments ICLR 2025 camera-ready version. Project page: https://tiger-ai-lab.github.io/MEGA-Bench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07839 2025-07-11 eess.IV cs.CV 79%

MeD-3D: A Multimodal Deep Learning Framework for Precise Recurrence Prediction in Clear Cell Renal Cell Carcinoma (ccRCC)

Hasaan Maqsood, Saif Ur Rehman Khan

机构 * Skolkovo Institute of Science and Technology (Skoltech)(斯克罗夫诺科学与技术学院(Skoltech)) German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07297 2025-07-11 cs.CV 79%

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning

Chengfei Wu, Ronald Seoh, Bingxuan Li, Liqiang Zhang, Fengrong Han, Dan Goldwasser

机构 * Purdue University(普渡大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California Los Angeles(加州大学洛杉矶分校) University of Connecticut(康涅狄格大学) University of California Berkeley(加州大学伯克利分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11150 2025-07-09 cs.CV cs.LG 79%

GC-GAT: Multimodal Vehicular Trajectory Prediction using Graph Goal Conditioning and Cross-context Attention

Mahir Gulzar, Yar Muhammad, Naveed Muhammad

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20167 2025-07-08 cs.CL 79%

Using Large Multimodal Models to Extract Knowledge Components for Knowledge Tracing from Multimedia Question Information

Hyeongdon Moon, Richard Davis, Seyed Parsa Neshaei, Pierre Dillenbourg

机构 * EPFL(瑞士联邦理工学院) KTH Royal Institute of Technology(皇家理工学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to Educational Data Mining 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10966 2025-07-08 eess.IV cs.CV 79%

ISLES'24: Final Infarct Prediction with Multimodal Imaging and Clinical Data. Where Do We Stand?

Ezequiel de la Rosa, Ruisheng Su, Mauricio Reyes, Evamaria O. Riedel, Hakim Baazaoui, Roland Wiest, Florian Kofler, Kaiyuan Yang, David Robben, Mahsa Mojtahedi, Laura van Poppel, Lucas de Vries, Anthony Winder, Kimberly Amador, Nils D. Forkert, Gyeongyeon Hwang, Jiwoo Song, Dohyun Kim, Eneko Uruñuela, Annabella Bregazzi, Matthias Wilms, Hyun Yang, Jin Tae Kwak, Sumin Jung, Luan Matheus Trindade Dalmazo, Kumaradevan Punithakumar, Moona Mazher, Abdul Qayyum, Steven Niederer, Jacob Idoko, Mariana Bento, Gouri Ginde, Tianyi Ren, Juampablo Heras Rivera, Mehmet Kurt, Carole Frindel, Susanne Wegener, Jan S. Kirschke, Benedikt Wiestler, Bjoern Menze

机构 * Department of Quantitative Biomedicine, University of Zurich, Zurich, Switzerland. Department of Biomedical Engineering, Eindhoven University of Technology, Eindhoven, the Netherlands. ARTORG Center for Biomedical Research, University of Bern, Bern, Switzerland. Department of Radiation Oncology, University Hospital Bern, University of Bern. Support Center of Advanced Neuroimaging (SCAN), University Institute of Diagnostic University Institute of Diagnostic Interventional Neuroradiology, University Hospital Bern, Inselspital, University of Bern, Bern, Switzerland. Department of Neurology, University Hospital of Zurich, Zurich, Switzerland. University of Zurich, Zurich, Switzerland. Department of Diagnostic Interventional Neuroradiology, School of Medicine Health, TUM Klinikum, Technical University of Munich, Germany. TranslaTUM, Center for Translational Cancer Research, Technical University of Munich, Germany. Therapy, School of Medicine Health, Technical University of Munich, Munich, Germany. Department of Biomedical Engineering Physics, Amsterdam UMC location University of Amsterdam, Amsterdam, the Netherlands Department of Radiology, University of Calgary, Calgary, Canada Hotchkiss Brain Institute, University of Calgary, Calgary, Canada Alberta Children’s Hospital Research Institute, University of Calgary, Calgary, Canada Heuron Co., Ltd., Seoul, South Korea School of Electrical Engineering, Korea University, Seoul, Korea University of Alberta, Alberta, Canada Hawkes Institute, Department of Computer Science, University College London, London, United Kingdom Lung Institute, Faculty of Medicine, Imperial College London, London, United Kingdom University of Calgary, Canada University of Washington, Washington, United States CREATIS, Universite Claude Bernard Lyon 1, INSA Lyon, UMR CNRS 5220, Inserm U1294, Villeurbanne, France Department of Neurology, University Hospital Zurich, Zurich, Switzerland

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03335 2025-07-08 cs.CL cs.CY 79%

iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to News

Tiancheng Hu, Nigel Collier

机构 * University of Cambridge(剑桥大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Dataset available at https://huggingface.co/datasets/pitehu/inews

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03013 2025-07-08 cs.CY cs.AI 79%

Challenges for AI in Multimodal STEM Assessments: a Human-AI Comparison

Aymeric de Chillaz, Anna Sotnikova, Patrick Jermann, Antoine Bosselut

机构 * EPFL(苏黎世联邦理工学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02932 2025-07-08 cs.LG cs.AI cs.CE 79%

MolProphecy: Bridging Medicinal Chemists' Knowledge and Molecular Pre-Trained Models via a Multi-Modal Framework

Jianping Zhao, Qiong Zhou, Tian Wang, Yusi Fan, Qian Yang, Li Jiao, Chang Liu, Zhehao Guo, Qi Lu, Fengfeng Zhou, Ruochi Zhang

机构 * College of Computer Science and Technology, Changchun University of Science and Technology(长春理工大学计算机科学与技术学院) Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology(香港科技大学化学与生物工程系) Key Laboratory of Symbolic Computation and Knowledge Engineering, Ministry of Education, Jilin University(吉林大学符号计算与知识工程重点实验室) Communication University of China(通信大学) Beijing Life Science Academy(北京生命科学研究院) College of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

Comments 16 pages,7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07416 2025-07-08 cs.CL 79%

ViMRHP: A Vietnamese Benchmark Dataset for Multimodal Review Helpfulness Prediction via Human-AI Collaborative Annotation

Truc Mai-Thanh Nguyen, Dat Minh Nguyen, Son T. Luu, Kiet Van Nguyen

机构 * Faculty of Information Science and Engineering(信息科学与工程学院) University of Information Technology(信息技术大学) Vietnam National University(越南国家大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at NLDB 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23918 2025-07-04 cs.CV 79%

Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Zhaochen Su, Peng Xia, Hangyu Guo, Zhenhua Liu, Yan Ma, Xiaoye Qu, Jiaqi Liu, Yanshu Li, Kaide Zeng, Zhengyuan Yang, Linjie Li, Yu Cheng, Heng Ji, Junxian He, Yi R. Fung

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) UNC-Chapel Hill(北卡罗来纳大学教堂山分校) Microsoft(微软) The Chinese University of Hong Kong(香港中文大学) UIUC(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Preprint in progress. We maintain a real-time GitHub repository tracking progress at: https://github.com/zhaochen0110/Awesome_Think_With_Images

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01045 2025-07-03 cs.LG cs.AI eess.SP 79%

Sensing Cardiac Health Across Scenarios and Devices: A Multi-Modal Foundation Model Pretrained on Heterogeneous Data from 1.7 Million Individuals

Xiao Gu, Wei Tang, Jinpei Han, Veer Sangha, Fenglin Liu, Shreyank N Gowda, Antonio H. Ribeiro, Patrick Schwab, Kim Branson, Lei Clifton, Antonio Luiz P. Ribeiro, Zhangdaihong Liu, David A. Clifton

机构 * University of Oxford(牛津大学) University of Nottingham(诺丁汉大学) Uppsala University(乌普萨拉大学) GlaxoSmithKline(葛兰素史克)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00490 2025-07-03 cs.CV eess.IV 79%

Just Noticeable Difference for Large Multimodal Models

Zijian Chen, Yuan Tian, Yuze Sun, Wei Sun, Zicheng Zhang, Weisi Lin, Guangtao Zhai, Wenjun Zhang

机构 * Institute of Image Communication and Information Processing, Shanghai Jiao Tong University, Shanghai 200240, China(上海交通大学图像通信与信息处理研究所) Shanghai AI Laboratory, Shanghai 200232, China(上海人工智能实验室) School of Communication and Electronic Engineering, East China Normal University, Shanghai 200241, China(华东师范大学通信与电子工程学院) School of Computer Science and Engineering, Nanyang Technological University, Singapore 639798(南洋理工大学计算机科学与工程学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 19 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16459 2025-07-03 cs.AI 79%

MMLU-Reason: Benchmarking Multi-Task Multi-modal Language Understanding and Reasoning

Guiyao Tie, Xueyang Zhou, Tianhe Gu, Ruihang Zhang, Chaoran Hu, Sizhe Zhang, Mengqu Sun, Yan Zhang, Pan Zhou, Lichao Sun

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

Comments 39 pages, 28 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23145 2025-07-01 cs.LG cs.CR cs.CV 79%

Forget-MI: Machine Unlearning for Forgetting Multimodal Information in Healthcare Settings

Shahad Hardan, Darya Taratynova, Abdelmajid Essofi, Karthik Nandakumar, Mohammad Yaqub

机构 * Department of Machine Learning, Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE Department of Computer Vision, Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE Department of Computer Science, Michigan State University, United States 0.2cm

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23122 2025-07-01 cs.CL cs.CY 79%

Decoding Memes: Benchmarking Narrative Role Classification across Multilingual and Multimodal Models

Shivam Sharma, Tanmoy Chakraborty

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15596 2025-07-01 cs.CV 79%

Mono-Modalizing Extremely Heterogeneous Multi-Modal Medical Image Registration

Kyobin Choo, Hyunkyung Han, Jinyeong Kim, Chanyong Yoon, Seong Jae Hwang

机构 * Department of Computer Science, Yonsei University(延世大学计算机科学系) Department of Artificial Intelligence, Yonsei University(延世大学人工智能系) Yonsei University College of Medicine(延世大学医学院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments 11 pages, 3 figures, 2 tables, Accepted at Medical Image Computing and Computer Assisted Intervention (MICCAI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21630 2025-06-30 cs.RO cs.CV cs.LG 79%

TOMD: A Trail-based Off-road Multimodal Dataset for Traversable Pathway Segmentation under Challenging Illumination Conditions

Yixin Sun, Li Li, Wenke E, Amir Atapour-Abarghouei, Toby P. Breckon

机构 * Department of Engineering, King’s College London, UK(工程系,伦敦国王学院,英国)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 9 figures, 2025 IJCNN

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24007 2025-06-30 cs.CV 79%

Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model

Nokimul Hasan Arif, Shadman Rabby, Md Hefzul Hossain Papon, Sabbir Ahmed

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Submitted for review in NCAA Springer, 21 pages, 4 figures, 4 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18104 2025-06-27 cs.MM 79%

Challenging Dataset and Multi-modal Gated Mixture of Experts Model for Remote Sensing Copy-Move Forgery Understanding

Ze Zhang, Enyuan Zhao, Yi Jiang, Jie Nie, Xinyue Liang

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.MM

Comments 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19615 2025-06-26 cs.CV 79%

Self-Supervised Multimodal NeRF for Autonomous Driving

Gaurav Sharma, Ravi Kothari, Josef Schmid

机构 * AVL Software and Functions GmbH(AVL软件与功能公司) Technische Hochschule Deggendorf(德格多夫技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23018 2025-06-26 cs.MM 79%

EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations

Haoqin Sun, Xuechen Wang, Jinghua Zhao, Shiwan Zhao, Jiaming Zhou, Hui Wang, Jiabei He, Aobo Kong, Xi Yang, Yequan Wang, Yonghua Lin, Yong Qin

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19860 2025-06-26 eess.SP cs.CV 79%

A Multi-Modal Spatial Risk Framework for EV Charging Infrastructure Using Remote Sensing

Oktay Karakuş, Padraig Corcoran

机构 * Cardiff University, School of Computer Science and Informatics(卡迪夫大学计算机科学与信息学学院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments 11 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19174 2025-06-25 cs.CV 79%

MOSCARD -- Causal Reasoning and De-confounding for Multimodal Opportunistic Screening of Cardiovascular Adverse Events

Jialu Pi, Juan Maria Farina, Rimita Lahiri, Jiwoong Jeong, Archana Gurudu, Hyung-Bok Park, Chieh-Ju Chao, Chadi Ayoub, Reza Arsanjani, Imon Banerjee

机构 * Department of Data Science & Engineering, Arizona State University(数据科学与工程系,亚利桑那州立大学) Department of Cardiovascular Medicine, Mayo Clinic(心血管医学系,梅奥诊所) Department of Radiology, Mayo Clinic(放射科,梅奥诊所) Catholic kwandong university international st. mary’s hospital(韩国天主教 kwandong 大学国际圣玛丽医院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18084 2025-06-24 cs.CV 79%

TEM^3-Learning: Time-Efficient Multimodal Multi-Task Learning for Advanced Assistive Driving

Wenzhuo Liu, Yicheng Qiao, Zhen Wang, Qiannan Guo, Zilong Chen, Meihua Zhou, Xinran Li, Letian Wang, Zhiwei Li, Huaping Liu, Wenshuo Wang

机构 * Faculty of Marine Science and Technology, Beijing Institute of Technology(海洋科学与技术学院,北京理工大学) State Key Laboratory of Intelligent Technology and Systems and Department of Computer Science and Technology, Tsinghua University(智能技术与系统国家重点实验室和计算机科学与技术系,清华大学) School of Engineering and Applied Science, Yale University(工程与应用科学学院,耶鲁大学) University of Toronto(多伦多大学) Beijing University of Chemical Technology(北京化工大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17958 2025-06-24 cs.CV 79%

ELMAR: Enhancing LiDAR Detection with 4D Radar Motion Awareness and Cross-modal Uncertainty

Xiangyuan Peng, Miao Tang, Huawei Sun, Bierzynski Kay, Lorenzo Servadei, Robert Wille

机构 * Infineon Technologies AG and the Technical University of Munich(英飞凌科技有限公司和技术大学慕尼黑) China University of Geosciences(中国地质大学) Technical University of Munich(技术大学慕尼黑)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments 7 pages. Accepted by IROS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17869 2025-06-24 cs.CV cs.RO 79%

Cross-modal State Space Modeling for Real-time RGB-thermal Wild Scene Semantic Segmentation

Xiaodong Guo, Zi'ang Lin, Luwen Hu, Zhihong Deng, Tong Liu, Wujie Zhou

机构 * School of Automation, Beijing Institute of Technology(自动化学院,北京理工大学) School of Information and Electronic Engineering, Zhejiang University of Science and Technology(信息电子工程学院,浙江工业大学) School of Computer Science and Engineering, Nanyang Technological University(计算机科学与工程学院,南洋理工大学)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17596 2025-06-24 cs.CV 79%

A Multimodal In Vitro Diagnostic Method for Parkinson's Disease Combining Facial Expressions and Behavioral Gait Data

Wei Huang, Yinxuan Xu, Yintao Zhou, Zhengyu Li, Jing Huang, Meng Pang

机构 * School of Mathematics and Computer Sciences, Nanchang University, Nanchang, China(南昌大学数学与计算机科学学院) Yichun University, Yichun, China(宜春大学) Nanchang University Second Affiliated Hospital, Nanchang, China(南昌大学第二附属医院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 4 figures, accepted by CogSci 2025

详情

展开后加载摘要…

URL PDF HTML 收藏