arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9189 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9189 篇

2505.16770 2025-05-26 cs.CV 79%

RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs

Meng-Hao Guo, Xuanyu Chu, Qianrui Yang, Zhe-Han Mo, Yiqing Shen, Pei-lin Li, Xinjie Lin, Jinnian Zhang, Xin-Sheng Chen, Yi Zhang, Kiyohiro Nakayama, Zhengyang Geng, Houwen Peng, Han Hu, Shi-Min Hu

机构 * Tsinghua University(清华大学) Tencent Hunyuan X(腾讯文英实验室) Stanford University(斯坦福大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10541 2025-05-26 cs.CV 79%

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis

Pengfei Wang, Guohai Xu, Weinong Wang, Junjie Yang, Jie Lou, Yunhua Xue

机构 * Pengfei Wang(王鹏飞) Guohai Xu(徐国海) Weinong Wang(王文龙) Junjie Yang(杨俊杰) Jie Lou(娄杰) Yunhua Xue(许云华)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01441 2025-05-26 cs.AI cs.LG 79%

LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations

Anian Ruoss, Fabio Pardo, Harris Chan, Bonnie Li, Volodymyr Mnih, Tim Genewein

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16365 2025-05-26 cs.CL 79%

Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines

Zi-Ao Ma, Tian Lan, Rong-Cheng Tu, Yong Hu, Yu-Shi Zhu, Tong Zhang, Heyan Huang, Zhijing Wu, Xian-Ling Mao

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17021 2025-05-23 cs.CV 79%

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark

Sara Ghaboura, Ketan More, Wafa Alghallabi, Omkar Thawakar, Jorma Laaksonen, Hisham Cholakkal, Salman Khan, Rao Muhammad Anwer

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Github : https://github.com/mbzuai-oryx/ARB, Huggingface: https://huggingface.co/datasets/MBZUAI/ARB

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15472 2025-05-23 cs.CL 79%

PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions

Song Dai, Yibo Yan, Jiamin Su, Dongfang Zihao, Yubo Gao, Yonghua Hei, Jungang Li, Junyan Zhang, Sicheng Tao, Zhuoran Gao, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Beijing Future Brain Education Technology Co., Ltd.(北京未来脑教育科技有限公司) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03644 2025-05-23 cs.CV 79%

DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms

Xiaojun Bi, Shuo Li, Junyao Xing, Ziyue Wang, Fuwen Luo, Weizheng Qiao, Lu Han, Ziwei Sun, Peng Li, Yang Liu

机构 * College of Information and Engineering, Minzu University of China(中国民族大学信息与工程学院) Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE, Minzu University of China(教育部民族语言智能分析与安全治理重点实验室) College of Information and Communication Engineering, Harbin Engineering University(哈尔滨工程大学信息与通信工程学院) Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University(清华大学人工智能研究院计算机科学与技术系) Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Our dataset can be obtained from: https://github.com/thinklis/DongbaMIE

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19140 2025-05-23 cs.CV 79%

Transformation trees -- documentation of multimodal image registration

Agnieszka Anna Tomaka, Dariusz Pojda, Michał Tarnawski, Leszek Luchowski

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 28 pages, 15 figures

Journal ref Computers in Biology and Medicine, Vol 193, year: 2025, pages: 110311

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15755 2025-05-22 cs.CV 79%

Exploring The Visual Feature Space for Multimodal Neural Decoding

Weihao Xia, Cengiz Oztireli

机构 * University of Cambridge(剑桥大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Project: https://weihaox.github.io/VINDEX

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15477 2025-05-22 cs.CV 79%

MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models

Mohammad Shahab Sepehri, Zalan Fabian, Maryam Soltanolkotabi, Mahdi Soltanolkotabi

机构 * Department of Electrical and Computer Engineering, University of Southern California(电气与计算机工程系,南加州大学) Department of Radiology and Imaging Sciences, University of Utah(放射学与影像科学系,犹他大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 24 Pages, 9 figures, The Thirteenth International Conference on Learning Representations (ICLR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02589 2025-05-21 cs.CL cs.IR 79%

MCiteBench: A Multimodal Benchmark for Generating Text with Citations

Caiyu Hu, Yikai Zhang, Tinghui Zhu, Yiwei Ye, Yanghua Xiao

机构 * Shanghai Key Laboratory of Data Science, School of Computer Science, Fudan University(上海数据科学 key 实验室,计算机科学学院,复旦大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments https://caiyuhu.github.io/MCiteBench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11936 2025-05-21 cs.CL 79%

A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges

Yibo Yan, Jiamin Su, Jianxiang He, Fangteng Fu, Xu Zheng, Yuanhuiyi Lyu, Kun Wang, Shen Wang, Qingsong Wen, Xuming Hu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by The 63rd Annual Meeting of the Association for Computational Linguistics (ACL Findings 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12766 2025-05-20 cs.CV 79%

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?

Haibin He, Maoyuan Ye, Jing Zhang, Xiantao Cai, Juhua Liu, Bo Du, Dacheng Tao

机构 * School of Computer Science, National Engineering Research Center for Multimedia Software, and Institute of Artificial Intelligence, Wuhan University, China(计算机学院、多媒体软件国家工程研究中心及人工智能研究所、武汉大学) College of Computing & Data Science at Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14321 2025-05-20 cs.CL 79%

Multimodal Coreference Resolution for Chinese Social Media Dialogues: Dataset and Benchmark Approach

Xingyu Li, Chen Gong, Guohong Fu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11962 2025-05-20 cs.AI 79%

CrafText Benchmark: Advancing Instruction Following in Complex Multimodal Open-Ended World

Zoya Volovikova, Gregory Gorbov, Petr Kuderov, Aleksandr I. Panov, Alexey Skrynnik

机构 * AIRI MIPT(莫斯科物理技术学院) FRC CSC RAS(俄罗斯科学院应用数学与控制论研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11907 2025-05-20 cs.CV 79%

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

Zihao Dongfang, Xu Zheng, Ziqiao Weng, Yuanhuiyi Lyu, Danda Pani Paudel, Luc Van Gool, Kailun Yang, Xuming Hu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09990 2025-05-20 cs.CV 79%

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

Long Cheng, Jiafei Duan, Yi Ru Wang, Haoquan Fang, Boyang Li, Yushan Huang, Elvis Wang, Ainaz Eftekhar, Jason Lee, Wentao Yuan, Rose Hendrix, Noah A. Smith, Fei Xia, Dieter Fox, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Allen Institute for Artificial Intelligence(人工智能 Allen 机构) Anderson Collegiate Vocational Institute(安德森职业学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 10 Pages, Dataset and code:https://pointarena.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14162 2025-05-19 cs.AI 79%

EIAD: Explainable Industrial Anomaly Detection Via Multi-Modal Large Language Models

Zongyun Zhang, Jiacheng Ruan, Xian Gao, Ting Liu, Yuzhuo Fu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10729 2025-05-19 eess.IV cs.CV q-bio.QM 79%

Adaptive Spatial Transcriptomics Interpolation via Cross-modal Cross-slice Modeling

NingFeng Que, Xiaofei Wang, Jingjing Chen, Yixuan Jiang, Chao Li

机构 * School of Science and Engineering, University of Dundee, UK(邓迪大学科学与工程学院) College of Medicine and Biological Information Engineering, Northeastern University, China(东北大学医学与生物信息工程学院) Department of Clinical Neurosciences, University of Cambridge, UK(剑桥大学临床神经科学系) School of Medicine, University of Dundee, UK(邓迪大学医学院) Department of Applied Mathematics and Theoretical Physics, University of Cambridge, UK(剑桥大学应用数学与理论物理系)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments Early accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24121 2025-05-19 cs.CV cs.LG 79%

IMPACT: A Generic Semantic Loss for Multimodal Medical Image Registration

Valentin Boussot, Cédric Hémon, Jean-Claude Nunes, Jason Dowling, Simon Rouzé, Caroline Lafond, Anaïs Barateau, Jean-Louis Dillenseger

机构 * Univ. Rennes, CLCC Eugene Marquis, INSERM, LTSI - UMR 1099(里昂大学、CLCC Eugene Marquis、INSERM、LTSI - UMR 1099) CSIRO Australian e-Health Research Centre(澳大利亚e-Health研究中心) CHU Rennes, Department of Cardio-Thoracic and Vascular Surgery(里昂大学医院、心胸和血管外科部门)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). This is a preprint version and has not been peer-reviewed

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09936 2025-05-16 cs.HC cs.GR cs.MA cs.MM 79%

CartoAgent: a multimodal large language model-powered multi-agent cartographic framework for map style transfer and evaluation

Chenglong Wang, Yuhao Kang, Zhaoya Gong, Pengjun Zhao, Yu Feng, Wenjia Zhang, Ge Li

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.MM

Comments 57 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09777 2025-05-16 cs.IR cs.CL 79%

A Survey on Large Language Models in Multimodal Recommender Systems

Alejo Lopez-Avila, Jinhua Du

机构 * Huawei London Research Centre(华为伦敦研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments 30 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08086 2025-05-14 cs.CV 79%

Multi-modal wound classification using wound image and location by Xception and Gaussian Mixture Recurrent Neural Network (GMRNN)

Ramin Mousa, Ehsan Matbooe, Hakimeh Khojasteh, Amirali Bengari, Mohammadmahdi Vahediahmar

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07396 2025-05-14 cs.CV cs.LG 79%

TUM2TWIN: Introducing the Large-Scale Multimodal Urban Digital Twin Benchmark Dataset

Olaf Wysocki, Benedikt Schwab, Manoj Kumar Biswanath, Michael Greza, Qilin Zhang, Jingwei Zhu, Thomas Froech, Medhini Heeramaglore, Ihab Hijazi, Khaoula Kanna, Mathias Pechinger, Zhaiyu Chen, Yao Sun, Alejandro Rueda Segura, Ziyang Xu, Omar AbdelGafar, Mansour Mehranfar, Chandan Yeshwanth, Yueh-Cheng Liu, Hadi Yazdi, Jiapan Wang, Stefan Auer, Katharina Anders, Klaus Bogenberger, Andre Borrmann, Angela Dai, Ludwig Hoegner, Christoph Holst, Thomas H. Kolbe, Ferdinand Ludwig, Matthias Nießner, Frank Petzold, Xiao Xiang Zhu, Boris Jutzi

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Submitted to the ISPRS Journal of Photogrammetry and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12133 2025-05-14 cs.HC cs.CV cs.LG 79%

VRMN-bD: A Multi-modal Natural Behavior Dataset of Immersive Human Fear Responses in VR Stand-up Interactive Games

He Zhang, Xinyang Li, Yuanxi Sun, Xinyi Fu, Christine Qiu, John M. Carroll

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to IEEE VR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07347 2025-05-13 cs.CV 79%

AI-Enabled Accurate Non-Invasive Assessment of Pulmonary Hypertension Progression via Multi-Modal Echocardiography

Jiewen Yang, Taoran Huang, Shangwei Ding, Xiaowei Xu, Qinhua Zhao, Yong Jiang, Jiarong Guo, Bin Pu, Jiexuan Zheng, Caojin Zhang, Hongwen Fei, Xiaomeng Li

机构 * Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(香港科技大学电子与计算机工程系) Guangdong Cardiovascular Institute, Guangdong Provincial People’s Hospital (Guangdong Academy of Medical Sciences), Southern Medical University(广东省心血管病研究所,广东省人民医院(广东省医学科学院)) Department of Ultrasound, The First Affiliated Hospital of Guangzhou Medical University(广州市第一人民医院超声科) Department of Pulmonary Circulation, Shanghai Pulmonary Hospital, Tongji University School of Medicine(上海 pulmonary 医院,同济大学医学院) Department of Echocardiography, Fuwai Hospital Chinese Academy of Medical Sciences(阜外医院中国医学科学院) Guangdong Provincial Key Laboratory of South China Structural Heart Disease(广东省南方结构性心脏病重点实验室) State Key Laboratory of Cardiovascular Disease, Department of Echocardiography, National Center for Cardiovascular Diseases, Fuwai Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College(心血管疾病国家重点实验室,国家心血管病中心,阜外医院,中国医学科学院和北京协和医学院) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06483 2025-05-13 cs.RO cs.CV 79%

CompSLAM: Complementary Hierarchical Multi-Modal Localization and Mapping for Robot Autonomy in Underground Environments

Shehryar Khattak, Timon Homberger, Lukas Bernreiter, Julian Nubert, Olov Andersson, Roland Siegwart, Kostas Alexis, Marco Hutter

机构 * Jet Propulsion Lab, California Institute of Technology, USA(喷气推进实验室,加州理工学院,美国) Division of Robotics, Perception, and Learning, KTH Royal Institute of Technology, Sweden(机器人、感知与学习系,皇家理工学院,瑞典) Autonomous Systems Lab, ETH Zürich, Switzerland(自主系统实验室,苏黎世联邦理工学院,瑞士) Robotic Systems Lab, ETH Zürich, Switzerland(机器人系统实验室,苏黎世联邦理工学院,瑞士) Autonomous Robots Lab, NTNU, Trondheim, Norway(自主机器人实验室,挪威特罗姆瑟大学,挪威)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments 8 pages, 9 figures, Code: https://github.com/leggedrobotics/compslam_subt

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05396 2025-05-13 cs.AI cs.HC cs.LG 79%

A Pain Assessment Framework based on multimodal data and Deep Machine Learning methods

Stefanos Gkikas

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14001 2025-05-12 cs.CV 79%

Multimodal Feature-Driven Deep Learning for the Prediction of Duck Body Dimensions and Weight

Wenbo Xiao, Qiannan Han, Gang Shu, Guiping Liang, Hongyan Zhang, Song Wang, Zhihao Xu, Weican Wan, Chuang Li, Guitao Jiang, Yi Xiao

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Journal ref Agriculture 2025, 15(10), 1021

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21699 2025-05-09 cs.CV cs.RO 79%

REHEARSE-3D: A Multi-modal Emulated Rain Dataset for 3D Point Cloud De-raining

Abu Mohammed Raisuddin, Jesper Holmblad, Hamed Haghighi, Yuri Poledna, Maikol Funk Drechsler, Valentina Donzella, Eren Erdal Aksoy

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏