arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9176 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9176 篇

2511.11623 2025-11-18 cs.LG cs.AI 79%

Early GVHD Prediction in Liver Transplantation via Multi-Modal Deep Learning on Imbalanced EHR Data

Yushan Jiang, Shuteng Niu, Dongjin Song, Yichen Wang, Jingna Feng, Xinyue Hu, Liu Yang, Cui Tao

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02650 2025-11-18 cs.CV 79%

Can Visual Input Be Compressed? A Visual Token Compression Benchmark for Large Multimodal Models

Tianfan Peng, Yuntao Du, Pengzhou Ji, Shijie Dong, Kailin Jiang, Mingchuan Ma, Yijun Tian, Jinhe Bi, Qian Li, Wei Du, Feng Xiao, Lizhen Cui

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09649 2025-11-18 cs.CV 79%

The Brain Resection Multimodal Image Registration (ReMIND2Reg) 2025 Challenge

Reuben Dorent, Laura Rigolo, Colin P. Galvin, Junyu Chen, Mattias P. Heinrich, Aaron Carass, Olivier Colliot, Demian Wassermann, Alexandra Golby, Tina Kapur, William Wells

机构 * Inria Saclay Île-de-France(法国里昂萨克利研究所) CEA(法国原子能委员会) Université Paris-Saclay(巴黎-萨克利大学) Sorbonne Université(索邦大学) Institut du Cerveau - Paris Brain Institute - ICM(巴黎脑研究所) CNRS(法国国家科学研究中心) Inria(法国国家信息与自动化技术研究所) Inserm(法国国家医学研究院) AP-HP(法国国家医院集团) Hôpital de la Pitié Salpêtrière(皮蒂埃-萨尔普里埃尔医院) Harvard Medical School(哈佛医学院) Brigham and Women's Hospital(布里根医院) Department of Radiology and Radiological Science(放射学与放射科学系) Johns Hopkins Medical School(约翰霍普金斯医学院) Institute of Medical Informatics(医学信息学研究所) University of Lübeck(吕贝克大学) Image Analysis and Communications Laboratory(图像分析与通信实验室) Department of Electrical and Computer Engineering(电气与计算机工程系) Johns Hopkins University(约翰霍普金斯大学) CSAIL(媒体实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11486 2025-11-17 cs.CV q-bio.QM 79%

Multimodal Posterior Sampling-based Uncertainty in PD-L1 Segmentation from H&E Images

Roman Kinakh, Gonzalo R. Ríos-Muñoz, Arrate Muñoz-Barrutia

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Preprint (pre-review). Accepted for publication in Lecture Notes in Bioinformatics (Springer, 2025). The final authenticated version will be available on SpringerLink once published

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11438 2025-11-17 cs.CV 79%

VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models

Mingjie Xu, Jinpeng Chen, Yuzhi Zhao, Jason Chun Lok Li, Yue Qiu, Zekang Du, Mengyang Wu, Pingping Zhang, Kun Li, Hongzheng Yang, Wenao Ma, Jiaheng Wei, Qinbin Li, Kangcheng Liu, Wenqiang Lei

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments This is the extended version of the paper accepted at AAAI 2026, which includes all technical appendices and additional experimental details

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11410 2025-11-17 cs.CV 79%

Q-Doc: Benchmarking Document Image Quality Assessment Capabilities in Multi-modal Large Language Models

Jiaxi Huang, Dongxu Wu, Hanwei Zhu, Lingyu Zhu, Jun Xing, Xu Wang, Baoliang Chen

机构 * South China Normal University(华南师范大学) Nanyang Technological University(南洋理工大学) City University of Hong Kong(香港城市大学) Shenzhen Academy of Inspection and Quarantine(深圳检验检疫局) Shenzhen University(深圳大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09914 2025-11-17 cs.AI 79%

OIDA-QA: A Multimodal Benchmark for Analyzing the Opioid Industry Documents Archive

Xuan Shen, Brian Wingenroth, Zichao Wang, Jason Kuen, Wanrong Zhu, Ruiyi Zhang, Yiwei Wang, Lichun Ma, Anqi Liu, Hongfu Liu, Tong Sun, Kevin S. Hawkins, Kate Tasker, G. Caleb Alexander, Jiuxiang Gu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by AAAI 2026 Artificial Intelligence for Social Impact Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15846 2025-11-17 cs.CL 79%

CyPortQA: Benchmarking Multimodal Large Language Models for Cyclone Preparedness in Port Operation

Chenchen Kuai, Chenhao Wu, Yang Zhou, Xiubin Bruce Wang, Tianbao Yang, Zhengzhong Tu, Zihao Li, Yunlong Zhang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09339 2025-11-13 cs.CL 79%

mmJEE-Eval: A Bilingual Multimodal Benchmark for Evaluating Scientific Reasoning in Vision-Language Models

Arka Mukherjee, Shreya Ghosh

机构 * Kalinga Institute of Industrial Technology (KIIT)(喀里亚理工学院) Indian Institute of Technology (IIT), Bhubaneswar(印度理工学院(班加罗尔))

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to IJCNLP-AACL Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08901 2025-11-13 cs.CV 79%

Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic Consistency

Riling Wei, Kelu Yao, Chuanguang Yang, Jin Wang, Zhuoyan Gao, Chao Li

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by AAAI-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22262 2025-11-12 cs.CV 79%

UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data

Yujian Yuan, Changjie Wu, Xinyuan Chang, Sijin Wang, Hang Zhang, Shiyi Liang, Shuang Zeng, Mu Xu, Ning Guo

机构 * Alibaba Group(阿里巴巴集团)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments AAAI2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13861 2025-11-12 cs.HC cs.CL cs.MA 79%

3MDBench: Medical Multimodal Multi-agent Dialogue Benchmark

Ivan Sviridov, Amina Miftakhova, Artemiy Tereshchenko, Galina Zubkova, Pavel Blinov, Andrey Savchenko

机构 * Sber AI Lab(Sber AI实验室) HSE University(俄罗斯高等经济大学) ISP RAS Research Center for Trusted Artificial Intelligence(俄罗斯科学院信息与系统问题研究所可信人工智能研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments EMNLP 25 (main)

Journal ref https://aclanthology.org/2025.emnlp-main.1353/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05883 2025-11-11 cs.AI 79%

Unveiling Modality Bias: Automated Sample-Specific Analysis for Multimodal Misinformation Benchmarks

Hehai Lin, Hui Liu, Shilei Cao, Jing Li, Haoliang Li, Wenya Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) City University of Hong Kong(香港城市大学) Sun Yat-sen University(中山大学) Harbin Institute of Technology(哈尔滨工业大学) Nanyang Technological University(新加坡国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18004 2025-11-11 cs.CV 79%

SkinCaRe: A Multimodal Dermatology Dataset Annotated with Medical Caption and Chain-of-Thought Reasoning

Yuhao Shen, Liyuan Sun, Yan Xu, Wenbin Liu, Shuping Zhang, Shawn Afvari, Zhongyi Han, Jiaoyan Song, Yongzhi Ji, Tao Lu, Xiaonan He, Xin Gao, Juexiao Zhou

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHK–Shenzhen)(数据科学学院,香港中文大学(深圳)) Computer Science Program, CEMSE Division, King Abdullah University of Science and Technology (KAUST)(计算机科学项目,科学与工程学院,国王 Abdullah 科学技术大学) Center of Excellence on Smart Health, KAUST(智能健康卓越中心,国王 Abdullah 科学技术大学) Center of Excellence for Generative AI, KAUST(生成式人工智能卓越中心,国王 Abdullah 科学技术大学) Department of Dermatology, Beijing AnZhen Hospital, Capital Medical University(皮肤科,北京安贞医院,首都医科大学) Department of Dermatology, Tianjin Institute of Integrative Dermatology, Tianjin Academy of Traditional Chinese Medicine Affiliated Hospital(皮肤科,天津整合皮肤科研究院,天津中医药大学附属医院) Department of Dermatology, Beijing Aerospace General Hospital(皮肤科,北京航天总医院) Department of Dermatology, The First Affiliated Hospital, Shantou University Medical College(皮肤科,汕头大学医学院第一附属医院) DermAssure, LLC(DermAssure 公司) School of Medicine, New York Medical College(医学院,纽约医学院) Capital Medical University(首都医科大学) Department of Dermatology, Second Hospital of Jilin University(皮肤科,吉林大学第二医院) Emergency Critical Care Center, Beijing AnZhen Hospital, Capital Medical University(急诊重症中心,北京安贞医院,首都医科大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04016 2025-11-07 cs.CV 79%

MedDChest: A Content-Aware Multimodal Foundational Vision Model for Thoracic Imaging

Mahmoud Soliman, Islam Osman, Mohamed S. Shehata, Rasika Rajapakshe

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03907 2025-11-07 cs.HC cs.AI 79%

SnappyMeal: Design and Longitudinal Evaluation of a Multimodal AI Food Logging Application

Liam Bakar, Zachary Englhardt, Vidya Srinivas, Girish Narayanswamy, Dilini Nissanka, Shwetak Patel, Vikram Iyer

机构 * University of Washington(华盛顿大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 24 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03768 2025-11-07 cs.LG cs.CV 79%

What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes

Candace Ross, Florian Bordes, Adina Williams, Polina Kirichenko, Mark Ibrahim

机构 * FAIR at Meta(Meta 的 FAIR)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 6 figures. Accepted to NeurIPS Datasets & Benchmarks 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.17168 2025-11-07 cs.CV 79%

EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUs

Zhen Fan, Peng Dai, Zhuo Su, Xu Gao, Zheng Lv, Jiarui Zhang, Tianyuan Du, Guidong Wang, Yang Zhang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18894 2025-11-06 cs.CV 79%

SmartWilds: Multimodal Wildlife Monitoring Dataset

Jenna Kline, Anirudh Potlapally, Bharath Pillai, Tanishka Wani, Rugved Katole, Vedant Patil, Penelope Covey, Hari Subramoni, Tanya Berger-Wolf, Christopher Stewart

机构 * Department of Computer Science and Engineering(计算机科学与工程系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to Imageomics Workshop at Neurips 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15748 2025-11-05 cs.AI 79%

Towards Relaxed Multimodal Inputs for Gait-based Parkinson's Disease Assessment

Minlin Zeng, Zhipeng Zhou, Yang Qiu, Martin J. McKeown, Zhiqi Shen

机构 * College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) Department of Medicine (Neurology), Pacific Parkinsons Research Centre, The University of British Columbia(医学系(神经学),太平洋帕金森研究中心,不列颠哥伦比亚大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22211 2025-11-05 cs.CL 79%

ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding

Kimihiro Hasegawa, Wiradee Imrattanatrai, Zhi-Qi Cheng, Masaki Asada, Susan Holm, Yuran Wang, Ken Fukuda, Teruko Mitamura

机构 * Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学) National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments NAACL2025, Code and Data: https://github.com/kimihiroh/promqa

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20241 2025-11-05 cs.LG cs.AI 79%

DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning

Qi Cao, Ruiyi Wang, Ruiyi Zhang, Sai Ashish Somayajula, Pengtao Xie

机构 * University of California, San Diego(加州大学圣地亚哥分校) Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 28 pages, 10 figures, to appear in NeurIPS 2025 (Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07215 2025-11-05 cs.RO cs.MM 79%

RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation

Feng Yan, Fanfan Liu, Liming Zheng, Yufeng Zhong, Yiyang Huang, Zechao Guan, Chengjian Feng, Lin Ma

机构 * Meituan(美团)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.MM

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01468 2025-11-04 cs.LG cs.AI 79%

DAMBench: A Multi-Modal Benchmark for Deep Learning-based Atmospheric Data Assimilation

Hao Wang, Zixuan Weng, Jindong Han, Wei Fan, Hao Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19954 2025-11-04 cs.AI cs.DB cs.LG 79%

RELATE: A Schema-Agnostic Perceiver Encoder for Multimodal Relational Graphs

Joe Meyer, Divyansha Lachi, Mahmoud Mohammadi, Roshan Reddy Upendra, Eva L. Dyer, Mark Li, Tom Palczewski

机构 * SAP Palo Alto, CA, USA(SAP帕洛阿尔托分校) University of Pennsylvania(宾夕法尼亚大学) SAP Seattle, WA, USA(SAP西雅图分校)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.AI

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01186 2025-11-04 cs.RO cs.CV 79%

LiDAR-VGGT: Cross-Modal Coarse-to-Fine Fusion for Globally Consistent and Metric-Scale Dense Mapping

Lijie Wang, Lianjie Guo, Ziyi Xu, Qianhao Wang, Fei Gao, Xieyuanli Chen

机构 * State Key Laboratory of Industrial Control Technology, Institute of Cyber-Systems and Control, Zhejiang University(工业控制技术国家重点实验室,系统与控制研究院,浙江大学) Differential Robot Technology Co., Ltd.(差分机器人技术有限公司) College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00584 2025-11-04 cs.IR cs.CL 79%

Structurally Refined Graph Transformer for Multimodal Recommendation

Ke Shi, Yan Zhang, Miao Zhang, Lifan Chen, Jiali Yi, Kui Xiao, Xiaoju Hou, Zhifei Li

机构 * School of Computer Science, Hubei University(湖北大学计算机学院) Hubei Key Laboratory of Big Data Intelligent Analysis and Application, Hubei University(湖北大学大数据智能分析与应用重点实验室) Key Laboratory of Intelligent Sensing System and Security (Hubei University), Ministry of Education(智能传感系统与安全重点实验室(湖北大学)) Institute of Vocational Education, Guangdong Industry Polytechnic University(广东职业技术学院教育学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments Comment: 13 pages, 7 figures, accepted by IEEE Transactions on Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00424 2025-11-04 cs.AI 79%

A Multimodal Framework for Depression Detection during Covid-19 via Harvesting Social Media: A Novel Dataset and Method

Ashutosh Anshul, Gumpili Sai Pranav, Mohammad Zia Ur Rehman, Nagendra Kumar

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Journal ref IEEE Transactions on Computational Social Systems, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00389 2025-11-04 cs.CV 79%

Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond

Fan Zhang, Haoxuan Li, Shengju Qian, Xin Wang, Zheng Lian, Hao Wu, Zhihong Zhu, Yuan Gao, Qiankun Li, Yefeng Zheng, Zhouchen Lin, Pheng-Ann Heng

机构 * The Chinese University of Hong Kong(香港中文大学) Peking University(北京大学) Tencent(腾讯) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学) Westlake University(西湖大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27481 2025-11-03 cs.CV 79%

NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding

Wei Xu, Cheng Wang, Dingkang Liang, Zongchuang Zhao, Xingyu Jiang, Peng Zhang, Xiang Bai

机构 * National University of Defense Technology(国防科技大学)

专题命中 多模态评测 :multimodal(title);image-text(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025. Data and models are available at https://github.com/H-EmbodVis/NAUTILUS

详情

展开后加载摘要…

URL PDF HTML 收藏