arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-26 至 2025-08-26 共收录 102 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 15 篇

2508.16859 2025-08-26 cs.CV 79%

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark

Jinpeng Hu, Hongchang Shi, Chongyuan Dai, Zhuo Li, Peipei Song, Meng Wang

机构 * Hefei University of Technology(合肥工业大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Science and Technology of China(中国科学技术大学) Institute of Artificial Intelligence (IAI), Hefei Comprehensive National Science Center(人工智能研究院(IAI),合肥综合性国家科学中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17880 2025-08-26 cs.HC 78%

TRUCE-AV: A Multimodal dataset for Trust and Comfort Estimation in Autonomous Vehicles

Aditi Bhalla, Christian Hellert, Enkelejda Kasneci, Nastassja Becker

专题命中 多模态评测 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17195 2025-08-26 astro-ph.SR astro-ph.HE 78%

A Multimodal Tiered Magnetic Polarity Inversion Features Dataset for Space Weather Forecasting

Ziba Khani, Anli Ji, Manolis K. Georgoulis, Berkay Aydin

专题命中 多模态评测 :multimodal(title,abstract)

Comments 21 pages, 10 figures. Submitted to Data in Brief (Elsevier)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14165 2025-08-26 eess.SP cs.SY eess.IV eess.SY 78%

A Multi-Modal IoT Node for Energy-Efficient Environmental Monitoring with Edge AI Processing

Philip Wiese, Victor Kartsch, Marco Guermandi, Luca Benini

专题命中 多模态评测 :multi-modal(title,abstract)

Comments 7 pages, 4 figures, 2 tables. This paper has been accepted at 2025 IEEE International Conference on Omni-layer Intelligent Systems (COINS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17290 2025-08-26 cs.AI cs.LG 74%

MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment

Omid Ghahroodi, Arshia Hemmat, Marzia Nouri, Seyed Mohammad Hadi Hosseini, Doratossadat Dastgheib, Mohammad Vali Sanian, Alireza Sahebi, Reihaneh Zohrabi, Mohammad Hossein Rohban, Ehsaneddin Asgari, Mahdieh Soleymani Baghshah

机构 * Computer Engineering Department, Sharif University of Technology, Iran(谢尔盖大学计算机工程系,伊朗) Qatar Computing Research Institute, Qatar(卡塔尔计算研究所,卡塔尔) Computer Engineering Department, University of Isfahan, Iran(伊斯法罕大学计算机工程系,伊朗) Independent Researcher(独立研究者)

专题命中 多模态评测 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16812 2025-08-26 cs.CV 74%

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes

Xinhao Xiang, Kuan-Chuan Peng, Suhas Lohit, Michael J. Jones, Jiawei Zhang

机构 * University of California, Davis(加州大学戴维斯分校) Mitsubishi Electric Research Laboratories (MERL)(三菱电机研究实验室)

专题命中 多模态评测 :multimodal(title);分类 cs.CV

Comments This paper is accepted to BMVC 2025 as an oral paper. The OVAD dataset is available at https://doi.org/10.5281/zenodo.16904069

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01608 2025-08-26 cs.CV cs.AI cs.LG q-bio.OT 62%

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models

Andy Bonnetto, Haozhe Qi, Franklin Leong, Matea Tashkovska, Mahdi Rad, Solaiman Shokur, Friedhelm Hummel, Silvestro Micera, Marc Pollefeys, Alexander Mathis

机构 * École Polytechnique Fédérale de Lausanne (EPFL)(瑞士联邦理工学院洛桑校区) Microsoft(微软) Scuola Superiore Sant’Anna(圣安娜高等学院) Swiss Federal Institute of Technology Valais (EPFL Valais)(瑞士联邦技术学院瓦莱分校) University of Geneva Medical School(日内瓦大学医学院) Eidgenössische Technische Hochschule (ETH)(瑞士联邦理工学院)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Code and data at: https://github.com/amathislab/EPFL-Smart-Kitchen

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16911 2025-08-26 cs.GR cs.CV cs.MM cs.SD 62%

MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation

Prerit Gupta, Jason Alexander Fotso-Puepi, Zhengyuan Li, Jay Mehta, Aniket Bera

机构 * Purdue University(普渡大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.MM

Comments Accepted at ICCV 2025. Project page: https://gprerit96.github.io/mdd-page

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16972 2025-08-26 cs.CV 57%

Robust Diagram Reasoning: A Framework for Enhancing LVLM Performance on Visually Perturbed Scientific Diagrams

Minghao Zhou, Rafael Souza, Yaqian Hu, Luming Che

机构 * Taiyuan University of Science and Technology(太原科技大学) University of Brasilia(巴西利亚大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16850 2025-08-26 cs.AI 57%

RADAR: A Reasoning-Guided Attribution Framework for Explainable Visual Data Analysis

Anku Rani, Aparna Garimella, Apoorv Saxena, Balaji Vasan Srinivasan, Paul Pu Liang

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 7 篇

2508.18108 2025-08-26 cs.CL 83%

SentiMM: A Multimodal Multi-Agent Framework for Sentiment Analysis in Social Media

Xilai Xu, Zilin Zhao, Chengye Song, Zining Wang, Jinhe Qiang, Jiongrui Yan, Yuhuai Lin

机构 * College of Information and Electrical Engineering, China Agricultural University(信息与电气工程学院,中国农业大学) College of Software, Jilin University(软件学院,吉林大学) College of Communication Engineering, Jilin University(通信工程学院,吉林大学)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00568 2025-08-26 cs.CV cs.AI 81%

Multimodal Masked Autoencoder Pre-training for 3D MRI-Based Brain Tumor Analysis with Missing Modalities

Lucas Robinet, Ahmad Berjaoui, Elizabeth Cohen-Jonathan Moyal

机构 * Oncopole(奥恩波勒) IRT Saint Exupéry(伊尔特圣埃克苏佩里研究所) INSERM Cancer Research Center of Toulouse(图卢兹癌症研究中心)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17398 2025-08-26 cs.CL 79%

DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards

Aaryaman Kartha, Ahmed Masry, Mohammed Saidul Islam, Thinh Lang, Shadikur Rahman, Ridwan Mahbub, Mizanur Rahman, Mahir Ahmed, Md Rizwan Parvez, Enamul Hoque, Shafiq Joty

机构 * York University(约克大学) RBC Qatar Computing Research Institute (QCRI)(卡塔尔计算研究所) Nanyang Technological University(南洋理工大学) Salesforce Research(Salesforce研究)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17205 2025-08-26 cs.CV cs.AI cs.CL eess.IV 67%

Multi-Agent Visual-Language Reasoning for Comprehensive Highway Scene Understanding

Yunxiang Yang, Ningning Xu, Jidong J. Yang

机构 * Smart Mobility and Infrastructure Lab(智能移动与基础设施实验室) College of Engineering, University of Georgia(佐治亚大学工程学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 16 pages, 16 figures, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06886 2025-08-26 cs.CV cs.AI cs.LG cs.MA cs.RO 62%

Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Yang Liu, Weixing Chen, Yongjie Bai, Xiaodan Liang, Guanbin Li, Wen Gao, Liang Lin

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) Sun Yat-sen University(中山大学) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室) Peng Cheng Laboratory(鹏城实验室) Institute of Digital Media, Peking University(北京大学数字媒体研究院)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

Comments The comprehensive review of Embodied AI. We also provide the resource repository for Embodied AI: https://github.com/HCPLab-SYSU/Embodied_AI_Paper_List

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17380 2025-08-26 cs.AI 57%

Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery

Jiaqi Liu, Songning Lai, Pengze Li, Di Yu, Wenjie Zhou, Yiyang Zhou, Peng Xia, Zijun Wang, Xi Chen, Shixiang Tang, Lei Bai, Wanli Ouyang, Mingyu Ding, Huaxiu Yao, Aoran Wang

机构 * UNC–Chapel Hill(北卡罗来纳大学教堂山分校) HKUST (Guangzhou)(香港科技大学(广州)) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Fudan University(复旦大学) Tsinghua University(清华大学) Nankai University(南开大学) UC Santa Cruz(圣塔克鲁兹大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17198 2025-08-26 cs.AI 57%

From reactive to cognitive: brain-inspired spatial intelligence for embodied agents

Shouwei Ruan, Liyuan Wang, Caixin Kang, Qihui Zhu, Songming Liu, Xingxing Wei, Hang Su

机构 * Department of Computer Science and Technology, Institute for AI, BNRist Center, Tsinghua-Bosch Joint ML Center, THBI Lab(计算机科学与技术系、人工智能研究院、BNRist中心、清华-博世联合机器学习中心、THBI实验室) Institute of Artificial Intelligence(人工智能研究院) Department of Psychological and Cognitive Sciences(心理学与认知科学系)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments 40 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 15 篇

2410.05849 2025-08-26 cs.CV 83%

ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt

Fanhu Zeng, Fei Zhu, Haiyang Guo, Xu-Yao Zhang, Cheng-Lin Liu

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(多模态人工智能系统国家重点实验室) School of Artificial Intelligence, UCAS(人工智能学院) Centre for Artificial Intelligence and Robotics, HKISI-CAS(人工智能与机器人中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16744 2025-08-26 cs.LG cs.CL cs.CV 81%

Hyperbolic Multimodal Representation Learning for Biological Taxonomies

ZeMing Gong, Chuanqi Tang, Xiaoliang Huo, Nicholas Pellegrino, Austin T. Wang, Graham W. Taylor, Angel X. Chang, Scott C. Lowe, Joakim Bruslund Haurum

机构 * Simon Fraser University(西蒙弗雷泽大学) University of Waterloo(滑铁卢大学) Vector Institute(向量研究所) University of Guelph(圭尔夫大学) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所) Aalborg University(奥胡斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18050 2025-08-26 cs.CV 79%

ArgusCogito: Chain-of-Thought for Cross-Modal Synergy and Omnidirectional Reasoning in Camouflaged Object Segmentation

Jianwen Tan, Huiyao Zhang, Rui Xiong, Han Zhou, Hongfei Wang, Ye Li

机构 * University of Chinese Academy of Sciences(中国科学院大学) Technology and Engineering Center for Space Utilization, Chinese Academy of Sciences(中国科学院空间利用技术与工程中心)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17478 2025-08-26 cs.CV 79%

GraphMMP: A Graph Neural Network Model with Mutual Information and Global Fusion for Multimodal Medical Prognosis

Xuhao Shan, Ruiquan Ge, Jikui Liu, Linglong Wu, Chi Zhang, Siqi Liu, Wenjian Qin, Wenwen Min, Ahmed Elazab, Changmiao Wang

机构 * Hangzhou Dianzi University(杭州电子科技大学) Shenzhen Polytechnic University(深圳职业技术大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Yunnan University(云南大学) Shenzhen University(深圳大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17213 2025-08-26 cs.CV 79%

Multi-modal Knowledge Decomposition based Online Distillation for Biomarker Prediction in Breast Cancer Histopathology

Qibin Zhang, Xinyu Hao, Qiao Chen, Rui Xu, Fengyu Cong, Cheng Lu, Hongming Xu

机构 * School of Biomedical Engineering, Faulty of Medicine, Dalian University of Technology, Dalian, China(生物医学工程学院) Faculty of Information Technology, University of Jyvaskyla, Jyvaskyla, Finland(信息科技学院) School of Software Technology, Dalian University of Technology, Dalian, China(软件技术学院) Department of Radiology, Guangdong Provincial People’s Hospital, Southern Medical University, Guangzhou, China(放射科)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17640 2025-08-26 eess.SP 78%

Multimodal Radio and Vision Fusion for Robust Localization in Urban V2I Communications

Can Zheng, Jiguang He, Chung G. Kang, Guofa Cai, Henk Wymeersch

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 6 pages, 6 figures, submitted to conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17524 2025-08-26 cs.CV cs.AI 73%

OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation

Xingxin He, Aurora Rofena, Ruimin Feng, Haozhe Liao, Zhaoye Zhou, Albert Jang, Fang Liu

机构 * Athinoula A. Martinos Center for Biomedical Imaging(阿提诺拉A.马丁诺斯生物医学成像中心) Harvard Medical School(哈佛医学院) Massachusetts General Hospital(麻省总医院) University Campus Bio-Medico of Rome(罗马生物医学大学校园)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17638 2025-08-26 cs.CV cs.CL 62%

Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning

Xinyu Wei, Guoli Yang, Jialu Zhou, Mingyue Yang, Leqian Li, Kedi Zhang, Chunping Qiu

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15123 2025-08-26 cs.CV cs.AI 62%

Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding

Ta Duc Huy, Duy Anh Huynh, Yutong Xie, Yuankai Qi, Qi Chen, Phi Le Nguyen, Sen Kim Tran, Son Lam Phung, Anton van den Hengel, Zhibin Liao, Minh-Son To, Johan W. Verjans, Vu Minh Hieu Phan

机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Macquarie University(麦考瑞大学) Hanoi University of Science and Technology(河内科学技术大学) University of Wollongong(沃林根大学) Flinders University(弗林德斯大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at ICCV 2025 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17619 2025-08-26 cs.CV 57%

Improving Interpretability in Alzheimer's Prediction via Joint Learning of ADAS-Cog Scores

Nur Amirah Abd Hamid, Mohd Shahrizal Rusli, Muhammad Thaqif Iman Mohd Taufek, Mohd Ibrahim Shapiai, Daphne Teck Ching Lai

机构 * School of Digital Science(数字科学学院) Universiti Brunei Darussalam(布鲁尼岛大学) Faculty of Artificial Intelligence(人工智能学院) Universiti Teknologi Malaysia(马来西亚理工大学) Malaysia-Japan International Institute of Technology(马来西亚-日本国际理工学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17283 2025-08-26 cs.CV cs.LG 57%

Quickly Tuning Foundation Models for Image Segmentation

Breenda Das, Lennart Purucker, Timur Carstensen, Frank Hutter

机构 * University of Freiburg(弗赖堡大学) ELLIS Institute Tübingen(图宾根ELLIS研究所) Prior Labs(Prior实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted as a short paper at the non-archival content track of AutoML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16974 2025-08-26 cs.CV 57%

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding

Leilei Guo, Antonio Carlos Rivera, Peiyu Tang, Haoxuan Ren, Zheyu Song

机构 * Zhongkai University of Agriculture and Engineering(仲恺农业工程大学) EDP University of Puerto Rico: San Sebastian(波多黎各圣塞巴斯蒂安EDP大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03430 2025-08-26 cs.LG cs.AI 57%

Multi-Level Fusion Graph Neural Network for Molecule Property Prediction

XiaYu Liu, Chao Fan, Yang Liu, Hou-biao Li

机构 * School of Mathematical Sciences, University of Electronic Science and Technology of China(电子科技大学数学科学学院) College of Management Science, Chengdu University of Technology(成都理工大学管理科学学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments 42 pages, 11 figures, 6 tables

Journal ref Journal of Chemical Information and Modeling,2025

详情

展开后加载摘要…

URL PDF HTML 收藏