arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9176 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9176 篇

2504.20466 2025-08-06 cs.CV 79%

LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMs

Woo Yi Yang, Jiarui Wang, Sijing Wu, Huiyu Duan, Yuxin Zhu, Liu Yang, Kang Fu, Guangtao Zhai, Xiongkuo Min

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01402 2025-08-05 cs.CV 79%

ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models

Chuangchuang Tan, Jinglu Wang, Xiang Ming, Renshuai Tao, Yunchao Wei, Yao Zhao, Yan Lu

机构 * Beijing Jiaotong University(北京交通大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21549 2025-08-05 cs.CV 79%

SiM3D: Single-instance Multiview Multimodal and Multisetup 3D Anomaly Detection Benchmark

Alex Costanzino, Pierluigi Zama Ramirez, Luigi Lella, Matteo Ragaglia, Alessandro Oliva, Giuseppe Lisanti, Luigi Di Stefano

机构 * CVLab, University of Bologna, Italy(博洛尼亚大学计算机视觉实验室) SACMI Imola, Italy(意大利SACMI公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2025. Project page: https://alex-costanzino.github.io/SiM3D/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13927 2025-08-05 cs.CV 79%

Multimodal 3D Reasoning Segmentation with Complex Scenes

Xueying Jiang, Lewei Lu, Ling Shao, Shijian Lu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09016 2025-08-05 cs.CV 79%

Cross-Modal Learning for Anomaly Detection in Complex Industrial Process: Methodology and Benchmark

Gaochang Wu, Yapeng Zhang, Lan Deng, Jingxin Zhang, Tianyou Chai

机构 * State Key Laboratory of Synthetical Automation for Process Industries, Northeastern University, Shenyang 110819, P. R. China(合成过程工业系统自动化国家重点实验室,东北大学,沈阳110819,中华人民共和国) School of Software and Electrical Engineering, Swinburne University of Technology, Melbourne, VIC 3122, Australia(软件与电子工程学院,斯威本科技大学,墨尔本,VIC 3122,澳大利亚)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments 14 pages, 6 figures, 5 tables. IEEE TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00726 2025-08-04 cs.CV 79%

MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models

Jiale Li, Mingrui Wu, Zixiang Jin, Hao Chen, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao, Rongrong Ji

机构 * Xiamen University(厦门大学) Zhongguancun Academy(中关村学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments ACM MM25 has accepted this paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14709 2025-08-04 cs.CV cs.RO 79%

DiFuse-Net: RGB and Dual-Pixel Depth Estimation using Window Bi-directional Parallax Attention and Cross-modal Transfer Learning

Kunal Swami, Debtanu Gupta, Amrit Kumar Muduli, Chirag Jaiswal, Pankaj Kumar Bajpai

机构 * Visual Intelligence Team, Samsung Research India Bangalore(三星印度班加罗尔视觉智能团队)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted in IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11172 2025-08-04 cs.CV 79%

TerraMesh: A Planetary Mosaic of Multimodal Earth Observation Data

Benedikt Blumenstiel, Paolo Fraccaro, Valerio Marsocci, Johannes Jakubik, Stefano Maurogiovanni, Mikolaj Czerkawski, Rocco Sedona, Gabriele Cavallaro, Thomas Brunschwiler, Juan Bernabe-Moreno, Nicolas Longépé

机构 * IBM Research – Europe(IBM欧洲研究院) European Space Agency(欧洲航天局) roman_Φ -Lab(Φ实验室) Forschungszentrum Jülich(尤利奇研究中心) University of Iceland(冰岛大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) Workshops

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20808 2025-08-04 cs.AI 79%

MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contexts

Peijie Wang, Zhong-Zhi Li, Fei Yin, Xin Yang, Dekang Ran, Cheng-Lin Liu

机构 * MAIS, Institute of Automation of Chinese Academy of Sciences(自动化研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 45 pages, accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23135 2025-08-01 cs.CL 79%

ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans

Ananya Sadana, Yash Kumar Lal, Jiawei Zhou

机构 * Stony Brook University(石溪大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14939 2025-08-01 cs.CV 79%

VisNumBench: Evaluating Number Sense of Multimodal Large Language Models

Tengjin Weng, Jingyi Wang, Wenhao Jiang, Zhong Ming

机构 * College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济发展实验室) Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际 Graduate School) Shenzhen Technology University(深圳技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22412 2025-07-31 cs.CV 79%

UAVScenes: A Multi-Modal Dataset for UAVs

Sijie Wang, Siqi Li, Yawei Zhang, Shangshu Yu, Shenghai Yuan, Rui She, Quanjiang Guo, JinXuan Zheng, Ong Kang Howe, Leonrich Chandra, Shrivarshann Srijeyan, Aditya Sivadas, Toshan Aggarwal, Heyuan Liu, Hongming Zhang, Chujie Chen, Junyu Jiang, Lihua Xie, Wee Peng Tay

机构 * Nanyang Technological University(南洋理工大学) School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院) Beihang University(北航大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21924 2025-07-30 cs.CV 79%

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning

Tianhong Gao, Yannian Fu, Weiqun Wu, Haixiao Yue, Shanshan Liu, Gang Zhang

机构 * Baidu Inc.(百度公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16876 2025-07-30 q-bio.QM cs.AI cs.LG 79%

Machine learning-based multimodal prognostic models integrating pathology images and high-throughput omic data for overall survival prediction in cancer: a systematic review

Charlotte Jennings, Andrew Broad, Lucy Godson, Emily Clarke, David Westhead, Darren Treanor

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Main article (50 pages, inc 3 tables, 4 figures). Supplementary material included with additional methodological information and data

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20764 2025-07-29 cs.CV 79%

ATR-UMMIM: A Benchmark Dataset for UAV-Based Multimodal Image Registration under Complex Imaging Conditions

Kangcheng Bin, Chen Chen, Ting Hu, Jiahao Qi, Ping Zhong

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20613 2025-07-29 cs.AI cs.LG 79%

Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression

Te Zhang, Yuheng Li, Junxiang Wang, Lujun Li

机构 * University of Michigan(密歇根大学) Johns Hopkins University(约翰霍普金斯大学) Central South university(中南大学) HKGAI

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20503 2025-07-29 cs.LG cs.CL cs.CY 79%

Customize Multi-modal RAI Guardrails with Precedent-based predictions

Cheng-Fu Yang, Thanh Tran, Christos Christodoulopoulos, Weitong Ruan, Rahul Gupta, Kai-Wei Chang

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11780 2025-07-29 cs.CV 79%

Rethinking Multi-Modal Object Detection from the Perspective of Mono-Modality Feature Learning

Tianyi Zhao, Boyang Liu, Yanglei Gao, Yiming Sun, Maoxun Yuan, Xingxing Wei

机构 * Institute of Artificial Intelligence, State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(人工智能研究院、虚拟现实技术与系统国家重点实验室、北京航空航天大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10568 2025-07-28 cs.CV 79%

AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark

Aruna Gauba, Irene Pi, Yunze Man, Ziqi Pang, Vikram S. Adve, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Rice University(Rice大学) Carnegie Mellon University(卡内基梅隆大学) AIFARMS Center for Digital Agriculture at UIUC(伊利诺伊大学厄巴纳-香槟分校数字农业中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Project Website: https://agmmu.github.io/ Huggingface: https://huggingface.co/datasets/AgMMU/AgMMU_v1/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17934 2025-07-25 cs.LG cs.AI 79%

Multimodal Fine-grained Reasoning for Post Quality Evaluation

Xiaoxu Guo, Siyan Liang, Yachao Cui, Juxiang Zhou, Lei Wang, Han Cao

机构 * School of Computer Science(计算机科学学院) School of Artificial Intelligence Institute(人工智能研究所) State Key Laboratory of Information Security, Institute of Information Engineering(信息安全国家重点实验室,信息工程研究所) Key Laboratory of Education Informatization for Nationalities(民族教育信息化重点实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 48 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16228 2025-07-23 cs.CV 79%

MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing

Shreelekha Revankar, Utkarsh Mall, Cheng Perng Phoo, Kavita Bala, Bharath Hariharan

机构 * Cornell University(康奈尔大学) Columbia University(哥伦比亚大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 17 pages, 9 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12545 2025-07-23 cs.CV 79%

PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models

Zhaopan Xu, Pengfei Zhou, Weidong Tang, Jiaxin Ai, Wangbo Zhao, Kai Wang, Xiaojiang Peng, Wenqi Shao, Hongxun Yao, Kaipeng Zhang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) School of Computer, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) School of Computer, National University of Singapore(计算机学院,国立新加坡大学) School of Computer, Xidian University(计算机学院,西安电子科技大学) College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳科技大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15255 2025-07-22 eess.SP cs.AI cs.LG 79%

MEETI: A Multimodal ECG Dataset from MIMIC-IV-ECG with Signals, Images, Features and Interpretations

Deyun Zhang, Xiang Lan, Shijia Geng, Qinghao Zhao, Sumei Fan, Mengling Feng, Shenda Hong

机构 * HeartVoice Medical Technology(HeartVoice医疗科技) Saw Swee Hock School of Public Health and Institute of Data Science(Saw Swee Hock公共卫生学院和数据科学研究所) National University of Singapore(新加坡国立大学) Department of Cardiology, Peking University People’s Hospital(北京大学人民医院心内科) College of Integrative Chinese and Western Medicine, Anhui University of Chinese Medicine(安徽中医药大学整合中西医学学院) National Institute of Health Data Science, Peking University(北京大学国家健康数据科学研究院) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01511 2025-07-22 cs.AI 79%

CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents

Tianqi Xu, Linyao Chen, Dai-Jie Wu, Yanjun Chen, Zecheng Zhang, Xiang Yao, Zhiqiang Xie, Yongchao Chen, Shilong Liu, Bochen Qian, Anjie Yang, Zhaoxuan Jin, Jianbo Deng, Philip Torr, Bernard Ghanem, Guohao Li

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 2025 ACL Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14824 2025-07-22 cs.LG cs.AI 79%

Benchmarking Foundation Models with Multimodal Public Electronic Health Records

Kunyu Yu, Rui Yang, Jingchi Liao, Siqi Li, Huitao Li, Irene Li, Yifan Peng, Rishikesan Kamaleswaran, Nan Liu

机构 * Centre for Quantitative Medicine and Duke-NUS AI + Medical Science Initiative, Duke-NUS Medical School(定量医学中心和杜克-国立新加坡大学AI+医学科学计划,杜克-国立新加坡大学医学院) Graduate School of Engineering, The University of Tokyo(东京大学工程研究生院) Department of Population Health Sciences, Weill Cornell Medicine(流行病学与公共卫生科学系,韦尔·科恩医学中心) Department of Surgery, Duke University School of Medicine(外科医学系,杜克大学医学学院) Centre for Quantitative Medicine, Duke-NUS AI + Medical Science Initiative and Programme in Health Services and Systems Research, Duke-NUS Medical School and NUS Artificial Intelligence Institute, National University of Singapore(定量医学中心和杜克-国立新加坡大学AI+医学科学计划及健康服务与系统研究计划,杜克-国立新加坡大学医学院和新加坡国立大学人工智能研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13812 2025-07-21 cs.CV 79%

SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing

Yingying Zhang, Lixiang Ru, Kang Wu, Lei Yu, Lei Liang, Yansheng Li, Jingdong Chen

机构 * Ant Group(蚂蚁集团) Wuhan University(武汉大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ICCV25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13318 2025-07-18 cs.CL 79%

HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals

Guimin Hu, Daniel Hershcovich, Hasti Seifi

机构 * University of Copenhagen(哥本哈根大学) Arizona State University(亚利桑那州立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22027 2025-07-17 cs.CV 79%

Cross-modal Ship Re-Identification via Optical and SAR Imagery: A Novel Dataset and Method

Han Wang, Shengyang Li, Jian Yang, Yuxuan Liu, Yixuan Lv, Zhuang Zhou

机构 * Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间利用技术与工程中心) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11114 2025-07-16 cs.CL 79%

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models

Seif Ahmed, Mohamed T. Younes, Abdelrahman Moustafa, Abdelrahman Allam, Hamza Moustafa

机构 * October University for Modern Sciences and Arts(现代科学与艺术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05763 2025-07-16 cs.LG cs.CL 79%

BMDetect: A Multimodal Deep Learning Framework for Comprehensive Biomedical Misconduct Detection

Yize Zhou, Jie Zhang, Meijie Wang, Lun Yu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏