arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9176 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9176 篇

2402.18319 2025-08-26 cs.RO cs.CV 79%

A Multimodal Handover Failure Detection Dataset and Baselines

Santosh Thoduka, Nico Hochgeschwender, Juergen Gall, Paul G. Plöger

机构 * Hochschule Bonn-Rhein-Sieg(波恩-莱茵-锡格应用科学大学) University of Bremen(不莱梅大学) University of Bonn(波恩大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔人工智能与机器学习研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICRA 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16859 2025-08-26 cs.CV 79%

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark

Jinpeng Hu, Hongchang Shi, Chongyuan Dai, Zhuo Li, Peipei Song, Meng Wang

机构 * Hefei University of Technology(合肥工业大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Science and Technology of China(中国科学技术大学) Institute of Artificial Intelligence (IAI), Hefei Comprehensive National Science Center(人工智能研究院(IAI),合肥综合性国家科学中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15299 2025-08-22 cs.CV 79%

BasketLiDAR: The First LiDAR-Camera Multimodal Dataset for Professional Basketball MOT

Ryunosuke Hayashi, Kohei Torimi, Rokuto Nagata, Kazuma Ikeda, Ozora Sako, Taichi Nakamura, Masaki Tani, Yoshimitsu Aoki, Kentaro Yoshioka

机构 * Keio University(Keio大学) AISIN CORPORATION(AISIN公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to MMSports

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13368 2025-08-21 cs.CV cs.LG 79%

MetaWild: A Multimodal Dataset for Animal Re-Identification with Environmental Metadata

Yuzhuo Li, Di Zhao, Tingrui Qiao, Yihao Wu, Bo Pang, Yun Sing Koh

机构 * University of Auckland(奥克兰大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 7 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07865 2025-08-21 cs.CV cs.RO 79%

AnoVox: A Benchmark for Multimodal Anomaly Detection in Autonomous Driving

Daniel Bogdoll, Iramm Hamdard, Lukas Namgyu Rößler, Felix Geisler, Muhammed Bayram, Felix Wang, Jan Imhof, Miguel de Campos, Anushervon Tabarov, Yitian Yang, Hanno Gottschalk, J. Marius Zöllner

机构 * FZI Research Center for Information Technology(弗劳恩霍夫研究所信息技术研究中心) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Technical University of Berlin(柏林技术大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Daniel Bogdoll, Iramm Hamdard, and Lukas Namgyu Rößler contributed equally. Accepted for publication at ECCV 2024 W-CODA workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.11370 2025-08-21 cs.CL 79%

G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Jiahui Gao, Renjie Pi, Jipeng Zhang, Jiacheng Ye, Wanjun Zhong, Yufei Wang, Lanqing Hong, Jianhua Han, Hang Xu, Zhenguo Li, Lingpeng Kong

机构 * Noah’s Ark Lab(诺亚 Ark 实验室) The University of Hong Kong(香港大学) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.CL

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14058 2025-08-21 cs.IR cs.AI 79%

Dual-Phase Playtime-guided Recommendation: Interest Intensity Exploration and Multimodal Random Walks

Jingmao Zhang, Zhiting Zhao, Yunqi Lin, Jianghong Ma, Tianjun Wei, Haijun Zhang, Xiaofeng Zhang

机构 * Harbin Institute of Technology(哈尔滨工业大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Accepted for publication at ACM Multimedia (ACM MM) 2025. 10 pages, 5 figures. Code and dataset: https://github.com/zqxwcevrtyui/DP2Rec

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06627 2025-08-20 cs.LG cs.AI 79%

Early Detection of Pancreatic Cancer Using Multimodal Learning on Electronic Health Records

Mosbah Aouad, Anirudh Choudhary, Awais Farooq, Steven Nevers, Lusine Demirkhanyan, Bhrandon Harris, Suguna Pappu, Christopher Gondi, Ravishankar Iyer

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Journal ref Proceedings of Machine Learning for Healthcare (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11654 2025-08-19 eess.SP cs.CV 79%

Data-driven RF Tomography via Cross-modal Sensing and Continual Learning

Yang Zhao, Tao Wang, Said Elhadi

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments 6 pages, 4 figures, to be published in IEEE AVSS Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04801 2025-08-15 cs.CV 79%

OrderChain: Towards General Instruct-Tuning for Stimulating the Ordinal Understanding Ability of MLLM

Jinhong Wang, Shuo Tong, Jian liu, Dongqi Tang, Weiqiang Wang, Wentong Li, Hongxia Xu, Danny Chen, Jintai Chen, Jian Wu

机构 * College of Computer Science & Technology, Zhejiang University(浙江大学计算机科学与技术学院) Transvascular Implantation Devices Research Institute and Liangzhu Laboratory(血管植入物研究机构和良渚实验室) Ant Group(蚂蚁集团) University of Notre Dame(圣母大学) HKUST (Guangzhou)(香港科技大学(广州))

专题命中 多模态评测 :MLLM(title);multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09715 2025-08-14 cs.CV cs.LG 79%

NEURAL: Attention-Guided Pruning for Unified Multimodal Resource-Constrained Clinical Evaluation

Devvrat Joshi, Islem Rekik

机构 * BASIRA Lab, Imperial-X (I-X) and Department of Computing, Imperial College London(BASIRA实验室、Imperial-X(I-X)及帝国理工学院计算机系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08071 2025-08-14 cs.LG cs.AI 79%

C-MAG: Cascade Multimodal Attributed Graphs for Supply Chain Link Prediction

Yunqing Li, Zixiang Tang, Jiaying Zhuang, Zhenyu Yang, Farhad Ameri, Jianbang Zhang

机构 * School of Manufacturing Systems and Networks, Arizona State University(制造系统与网络学院,亚利桑那州立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments https://openreview.net/pdf?id=mE5n6OJHwO

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19256 2025-08-14 cs.CV cs.RO 79%

LM-MCVT: A Lightweight Multi-modal Multi-view Convolutional-Vision Transformer Approach for 3D Object Recognition

Songsong Xiong, Hamidreza Kasaei

机构 * Department of Artificial Intelligence, University of Groningen(人工智能系,格罗宁根大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09155 2025-08-14 cs.LG cs.AI 79%

A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal Models

Wenkai Wang, Hongcan Guo, Zheqi Lv, Shengyu Zhang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 17 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01306 2025-08-13 cs.LG cs.CV 79%

ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation

Moran Yanuka, Morris Alper, Hadar Averbuch-Elor, Raja Giryes

机构 * Tel-Aviv University(特拉维夫大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to ACL 2024 (Finding). For Project webpage, see https://moranyanuka.github.io/icc/

Journal ref Findings of the Association for Computational Linguistics: ACL 2024, pages 11048-11064, Bangkok, Thailand, August 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07628 2025-08-12 cs.AI 79%

Multimodal AI Systems for Enhanced Laying Hen Welfare Assessment and Productivity Optimization

Daniel Essien, Suresh Neethirajan

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 66 pages, 7 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07596 2025-08-12 cs.CV 79%

From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users

Shahroz Tariq, Simon S. Woo, Priyanka Singh, Irena Irmalasari, Saakshi Gupta, Dev Gupta

机构 * Sungkyunkwan University, S. Korea(顺天大学) University of Queensland, Australia(昆士兰大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 3 tables, 5 figures, accepted for publicaiton in the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06851 2025-08-12 cs.AI cs.CY 79%

MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams

Pengfei Zhou, Xiaopeng Peng, Fanrui Zhang, Zhaopan Xu, Jiaxin Ai, Yansheng Qiu, Chuanhao Li, Zhen Li, Ming Li, Yukang Feng, Jianwen Sun, Haoquan Zhang, Zizhen Li, Xiaofeng Mao, Zekai Li, Wangbo Zhao, Kai Wang, Xiaojun Chang, Wenqi Shao, Yang You, Kaipeng Zhang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) USTC(中国科学技术大学) RIT(罗切斯特理工学院) HIT(哈尔滨工业大学) WHU(武汉大学) MBZUAI(马克斯·普朗克人工智能研究所) NUS(新加坡国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 35 pages, 33 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06558 2025-08-12 cs.CV cs.LG 79%

On the effectiveness of multimodal privileged knowledge distillation in two vision transformer based diagnostic applications

Simon Baur, Alexandra Benova, Emilio Dolgener Cantú, Jackie Ma

机构 * Fraunhofer Heinrich-Hertz-Institut(弗劳恩霍夫 Heinrich-Hertz 研究所) Universität Osnabrück(奥斯纳布吕克大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06009 2025-08-11 cs.CV 79%

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models

Jun Feng, Zixin Wang, Zhentao Zhang, Yue Guo, Zhihan Zhou, Xiuyi Chen, Zhenyang Li, Dawei Yin

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 29 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05527 2025-08-08 cs.CV 79%

AI vs. Human Moderators: A Comparative Evaluation of Multimodal LLMs in Content Moderation for Brand Safety

Adi Levi, Or Levi, Sardhendu Mishra, Jonathan Morra

机构 * Zefr Inc(Zefr公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to the Computer Vision in Advertising and Marketing (CVAM) workshop at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05083 2025-08-08 cs.AI 79%

MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models

Dexuan Xu, Jieyi Wang, Zhongyan Chai, Yongzhi Cao, Hanpin Wang, Huamin Zhang, Yu Huang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05053 2025-08-08 cs.CV 79%

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?

Parth Thakkar, Ankush Agarwal, Prasad Kasu, Pulkit Bansal, Chaitanya Devaguptapu

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Accepted at ACL 2025 in the main track

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04088 2025-08-08 cs.CL 79%

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning

Jianghangfan Zhang, Yibo Yan, Kening Zheng, Xin Zou, Song Dai, Xuming Hu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04192 2025-08-07 cs.CV 79%

From Learning to Unlearning: Biomedical Security Protection in Multimodal Large Language Models

Dunyuan Xu, Xikai Yang, Yaoqian Li, Jinpeng Li, Pheng-Ann Heng

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04059 2025-08-07 cs.CV 79%

Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models

Zhaochen Liu, Kaiwen Gao, Shuyi Liang, Bin Xiao, Limeng Qiao, Lin Ma, Tingting Jiang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04017 2025-08-07 cs.CV 79%

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability

Haiqi Yang, Jinzhe Li, Gengxu Li, Yi Chang, Yuan Wu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 9pages, 2figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03173 2025-08-06 cs.AI 79%

Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary Constructions

Jingxuan Wei, Caijun Jia, Qi Chen, Honghao He, Linzhuang Sun, Conghui He, Lijun Wu, Bihui Yu, Cheng Tan

机构 * Shenyang institute of computing technology, Chinese academy of sciences(沈阳计算技术研究所,中国科学院) Shanghai AI Laboratory(上海人工智能实验室) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02957 2025-08-06 eess.IV cs.CV 79%

AMD-Mamba: A Phenotype-Aware Multi-Modal Framework for Robust AMD Prognosis

Puzhen Wu, Mingquan Lin, Qingyu Chen, Emily Y. Chew, Zhiyong Lu, Yifan Peng, Hexin Dong

机构 * Department of Population Health Sciences, Weill Cornell Medicine(人口健康科学系,韦尔·柯林斯医学中心) Department of Surgery, University of Minnesota(外科系,明尼苏达大学) Department of Biomedical Informatics and Data Science, Yale School of Medicine(生物医学信息学与数据科学系,耶鲁医学院) National Eye Institute, National Institutes of Health(国家眼科研究所,国立卫生研究院) National Library of Medicine, National Institutes of Health(国家医学图书馆,国立卫生研究院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at the MICCAI 2025 MIML Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02951 2025-08-06 cs.AI 79%

MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine

Mahtab Bigverdi, Wisdom Ikezogwo, Kevin Zhang, Hyewon Jeong, Mingyu Lu, Sungjae Cho, Linda Shapiro, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Massachusetts Institute of Technology(麻省理工学院) Seoul National University(首尔国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏