arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-19 至 2025-11-19 共收录 56 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 12 篇

2412.15925 2025-11-19 cs.CV cs.AI 84%

MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Classification and Detection

Andrea Moglia, Elia Clement Nastasio, Luca Mainardi, Pietro Cerveri

机构 * Department of Electronics, Information, and Bioengineering(电子、信息与生物工程系) Polytechnic University of Milan(米兰理工学院) Department of Industrial, and Information Engineering(工业与信息工程系) University of Pavia(帕维亚大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Journal ref Moglia, A., Nastasio, E.C., Mainardi, L. et al. MiniGPT-Pancreas: Multimodal Large Language Model for Pancreas Cancer Observation and Localization in CT Images. J Healthc Inform Res (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23947 2025-11-19 cs.HC 82%

The Social Gaze of LLMs: A Literature Review of Multimodal Approaches to Human Behavior Understanding

Zihan Liu, Parisa Rabbani, Veda Duddu, Kyle Fan, Madison Lee, Yun Huang

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11777 2025-11-19 cs.RO cs.CV 70%

Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy

Vinit Mehta, Charu Sharma, Karthick Thiyagarajan

机构 * Machine Learning Lab IIIT Hyderabad(IIIT Hyderabad 机器学习实验室) SensR Lab Western Sydney University(Western Sydney University SensR实验室)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 45 pages, 15 figures, MDPI Sensors Journal

Journal ref Sensors 2025, 25(20), 6394

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22995 2025-11-19 cs.CV cs.AI cs.CL cs.LG 67%

VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning

Jingkun Ma, Runzhe Zhan, Yang Li, Di Sun, Hou Pong Chan, Lidia S. Chao, Derek F. Wong

机构 * NLP(自然语言处理) CT Lab, Department of Computer and Information Science, University of Macau(计算机与信息科学系计算机视觉实验室,澳门大学) University of Macau(澳门大学) DAMO Academy, Alibaba Group(阿里集团达摩院)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 58 pages, 28 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14082 2025-11-19 cs.CV cs.AI 62%

Zero-Training Task-Specific Model Synthesis for Few-Shot Medical Image Classification

Yao Qin, Yangyang Yan, YuanChao Yang, Jinhua Pang, Huanyong Bi, Yuan Liu, HaiHua Wang

机构 * AI Innovation Department, Beijing 1st BioTech Group Co., Ltd.(人工智能创新部门,北京第一生物科技集团有限公司) Diplomatic Negotiation Simulation and Data Laboratory, China Foreign Affairs University(外交谈判模拟与数据实验室,中国外交学院)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14131 2025-11-19 cs.AI 57%

Run, Ruminate, and Regulate: A Dual-process Thinking System for Vision-and-Language Navigation

Yu Zhong, Zihao Zhang, Rui Zhang, Lingdong Huang, Haihan Gao, Shuo Wang, Da Li, Ruijian Han, Jiaming Guo, Shaohui Peng, Di Huang, Yunji Chen

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13909 2025-11-19 cs.CV 57%

Mind the Gap: Evaluating LLM Understanding of Human-Taught Road Safety Principles

Chalamalasetti Kranti

机构 * uni-potsdam(波恩大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13892 2025-11-19 cs.AI 57%

Jailbreaking Large Vision Language Models in Intelligent Transportation Systems

Badhan Chandra Das, Md Tasnim Jawad, Md Jueal Mia, M. Hadi Amini, Yanzhao Wu

机构 * KFSCIS, Florida International University(凯斯-西储大学信息科学学院) Knight Foundation School of Computing and Information Sciences(骑士基金会计算与信息科学学院) Learning for InterDependent Networks Laboratory (solid lab)(依赖网络学习实验室)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10094 2025-11-19 cs.LG cs.CV 57%

How does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders

Yiming Tang, Abhijeet Sinha, Dianbo Liu

机构 * National University of Singapore(新加坡国立大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14359 2025-11-19 cs.HC cs.SE 50%

Towards LLM-Based Usability Analysis for Recommender User Interfaces

Sebastian Lubos, Alexander Felfernig, Damian Garber, Viet-Man Le, Thi Ngoc Trang Tran

专题命中 多模态评测 :multimodal(abstract)

Comments The paper was presented at IntRS'25: Joint Workshop on Interfaces and Human Decision Making for Recommender Systems, September 22, 2025, Prague, Czech Republic and is published in the workshop proceedings: https://ceur-ws.org/Vol-4027/

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00781 2025-11-19 cs.LG stat.ML 50%

Structured Radial Basis Function Network: Modelling Diversity for Multiple Hypotheses Prediction

Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong

机构 * Department of Computer Science, University of Reading, United Kingdom(计算机科学系,阅读大学,英国)

专题命中 多模态评测 :multi-modal(abstract)

Comments Acepted Paper for AI-2024 Forty-fourth SGAI International Conference on Artificial Intelligence CAMBRIDGE, ENGLAND 17-19 DECEMBER 2024

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 1 篇

2508.14160 2025-11-19 cs.CV cs.AI cs.RO 66%

RynnEC: Bringing MLLMs into Embodied World

Ronghao Dang, Yuqian Yuan, Yunxuan Mao, Kehan Li, Jiangpin Liu, Zhikai Wang, Xin Li, Fan Wang, Deli Zhao

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab(虎盘实验室) Zhejiang University(浙江大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI;MLLM(comments)

Comments The technical report of RynnEC, an embodied cognition MLLM

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 9 篇

2511.14604 2025-11-19 cs.CV 83%

XAttn-BMD: Multimodal Deep Learning with Cross-Attention for Femoral Neck Bone Mineral Density Estimation

Yilin Zhang, Leo D. Westbury, Elaine M. Dennison, Nicholas C. Harvey, Nicholas R. Fuggle, Rahman Attar

机构 * School of Electronics and Computer Science, University of Southampton, UK(电子与计算机科学学院,索姆塞特大学,英国) MRC Lifecourse Epidemiology Centre, University of Southampton, Southampton General Hospital, UK(生命课程流行病学研究中心,索姆塞特大学,南安普顿总医院,英国) NIHR Southampton Biomedical Research Centre, University of Southampton(南安普顿生物医学研究中心,索姆塞特大学) University Hospital NHS Foundation Trust, Southampton, UK(南安普顿国家健康服务基金会信托,英国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 11 figures, 10 tables, 38 pages. Submitted to Artificial Intelligence in Medicine (currently with editor)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13755 2025-11-19 cs.LG cs.AI 83%

Adaptive Redundancy Regulation for Balanced Multimodal Information Refinement

Zhe Yang, Wenrui Li, Hongtao Chen, Penghong Wang, Ruiqin Xiong, Xiaopeng Fan

机构 * Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Harbin Institute of Technology Zhengzhou Research Institute(哈尔滨工业大学郑州研究所) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究所) School of Mathematical Sciences, University of Electronic Science and Technology of China(数学学院,电子科学与技术大学) School of Electronic Engineering and Computer Science, Institute of Digital Media, Peking University(电子工程与计算机科学系,数字媒体研究所,北京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13794 2025-11-19 cs.CV cs.AI 81%

FusionFM: All-in-One Multi-Modal Image Fusion with Flow Matching

Huayi Zhu, Xiu Shu, Youqiang Xiong, Qiao Liu, Rui Chen, Di Yuan, Xiaojun Chang, Zhenyu He

机构 * Guangzhou Institute of Technology, Xidian University(广州理工大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14693 2025-11-19 cs.CL 79%

Talk, Snap, Complain: Validation-Aware Multimodal Expert Framework for Fine-Grained Customer Grievances

Rishu Kumar Singh, Navneet Shreya, Sarmistha Das, Apoorva Singh, Sriparna Saha

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments To be published in the Proceedings of the 40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026 Special Track on AI for Social Impact )

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14157 2025-11-19 cs.CV 79%

Learning Representation and Synergy Invariances: A Povable Framework for Generalized Multimodal Face Anti-Spoofing

Xun Lin, Shuai Wang, Yi Yu, Zitong Yu, Jiale Zhou, Yizhong Liu, Xiaochun Cao, Alex Kot, Yefeng Zheng

机构 * Beihang University(北京航空航天大学) Nanyang Technological University(南洋理工大学) Westlake University(西湖大学) Great Bay University(大亚湾大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15127 2025-11-19 cs.LG 79%

PRIMUS: Pretraining IMU Encoders with Multimodal Self-Supervision

Arnav M. Das, Chi Ian Tang, Fahim Kawsar, Mohammad Malekzadeh

机构 * Nokia Bell Labs Cambridge, UK(诺基亚贝尔实验室(剑桥,英国)) University of Washington, USA(华盛顿大学(美国)) University of Glasgow, UK(格拉斯哥大学(英国))

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments Presented at ICASSP 2025. Also presented under the title "PRIMUS: Pretraining IMU Encoders with Multimodal and Self-Supervised Learning" at NeurIPS 2024 TSALM Workshop (Time Series in the Age of Large Models)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14601 2025-11-19 cs.CV cs.AI 62%

MRI Embeddings Complement Clinical Predictors for Cognitive Decline Modeling in Alzheimer's Disease Cohorts

Nathaniel Putera, Daniel Vilet Rodríguez, Noah Videcrantz, Julia Machnio, Mostafa Mehdipour Ghazi

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted at SPIE - Medical Imaging Conference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14698 2025-11-19 cs.CV cs.LG eess.SP 57%

HyMAD: A Hybrid Multi-Activity Detection Approach for Border Surveillance and Monitoring

Sriram Srinivasan, Srinivasan Aruchamy, Siva Ram Krisha Vadali

机构 * Sriram Srinivasan(独立研究者) Srinivasan Aruchamy(独立研究者) Siva Ram Krishna Vadali(独立研究者)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Multi-label seismic signal classification using novel attention-based feature fusion. Submitting to cs.CV due to relevance to general pattern recognition and time-frequency (spectrogram) analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14064 2025-11-19 cs.LG cs.AI stat.ME 57%

CafeMed: Causal Attention Fusion Enhanced Medication Recommendation

Kelin Ren, Chan-Yang Ju, Dong-Ho Lee

机构 * Department of Computer Science and Engineering, Hanyang University(韩阳大学计算机科学与工程系) Department of Applied Artificial Intelligence, Hanyang University(韩阳大学应用人工智能系)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.AI

Comments Accepted by BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 5 篇

2504.04681 2025-11-19 stat.ME 71%

Multimodal Distributions for Circular Axial Data

Fernández-Durán, J. J., Gregorio-Domínguez, M. M

专题命中 其他多模态 :multimodal(title)

Comments 28 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14341 2025-11-19 cs.RO cs.AI cs.CV 62%

Going Places: Place Recognition in Artificial and Natural Systems

Michael Milford, Tobias Fischer

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Journal ref Annual Review of Control, Robotics, and Autonomous Systems 2026, vol. 9

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14669 2025-11-19 q-bio.MN 50%

Hyperbolic Graph Embeddings Reveal the Host-Pathogen Interactome

Xiaoqiong Xia, Cesar de la Fuente-Nunez

专题命中 其他多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14517 2025-11-19 eess.SP 50%

Tri-Hybrid Beamforming Design for Fully-Connected Pinching Antenna Systems

Cheng-Jie Zhao, Zhaolin Wang, Hyundong Shin, Yuanwei Liu

专题命中 其他多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23578 2025-11-19 physics.flu-dyn 50%

Competing Mechanisms at Vibrated Interfaces of Density-Contrast Fluids

Tianyi Chu, Benjamin Wilfong, Timothy Koehler, Ryan M. McMullen, Spencer H. Bryngelson

专题命中 其他多模态 :multi-modal(abstract)

Journal ref Physical Review Fluids, 10, 093904 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏