arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-23 至 2025-09-23 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 16 篇

2504.04653 2025-09-23 cs.CV cs.CL 88%

LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts

Yimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof Czarnecki

机构 * University of Waterloo(滑铁卢大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(title);MLLM(abstract);分类 cs.CV、cs.CL

Comments To appear at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17136 2025-09-23 cs.CV cs.AI 84%

SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM

Yuhao Tian, Zheming Yang

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of AI for Industries(工业人工智能研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16900 2025-09-23 cs.CV cs.AI 84%

ME-Mamba: Multi-Expert Mamba with Efficient Knowledge Capture and Fusion for Multimodal Survival Analysis

Chengsheng Zhang, Linhao Qu, Xiaoyu Liu, Zhijian Song

机构 * Digital Medical Research Center, School of Basic Medical Science, Fudan University, Shanghai 200032, China(复旦大学基础医学学院数字医学研究中心) Shanghai Key Lab of Medical Image Computing(上海医学图像计算重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16618 2025-09-23 cs.CV cs.AI 84%

Surgical-MambaLLM: Mamba2-enhanced Multimodal Large Language Model for VQLA in Robotic Surgery

Pengfei Hao, Hongqiu Wang, Shuaibo Li, Zhaohu Xing, Guang Yang, Kaishun Wu, Lei Zhu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Imperial College London(帝国理工学院伦敦分校) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Early accepted by MICCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21059 2025-09-23 cs.CV cs.AI cs.CR cs.LG 81%

FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts

Ziyi Zhang, Zhen Sun, Zongmin Zhang, Jihui Guo, Xinlei He

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17943 2025-09-23 cs.CV cs.LG 79%

Can multimodal representation learning by alignment preserve modality-specific information?

Romain Thoreau, Jessie Levillain, Dawa Derksen

机构 * institutetext(机构文本) CNES(法国国家空间研究中心) INSA-IMT(法国里尔INSA-IMT)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted as a workshop paper at MACLEAN - ECML/PKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17520 2025-09-23 cs.CV 79%

Unified Multimodal Coherent Field: Synchronous Semantic-Spatial-Vision Fusion for Brain Tumor Segmentation

Mingda Zhang, Yuyang Zheng, Ruixiang Tang, Jingru Qiu, Haiyan Ding

机构 * School of Software, Yunnan University(云南大学软件学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17228 2025-09-23 cs.LG cs.CL stat.ME 79%

Causal Representation Learning from Multimodal Clinical Records under Non-Random Modality Missingness

Zihan Liang, Ziwen Pan, Ruoxuan Xiong

机构 * Emory University(埃默里大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments To appear in Proc. of EMNLP 2025 (18 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18005 2025-09-23 cs.RO 78%

M3ET: Efficient Vision-Language Learning for Robotics based on Multimodal Mamba-Enhanced Transformer

Yanxin Zhang, Liang He, Zeyi Kang, Zuheng Ming, Kaixing Zhao

机构 * School of Software Northwestern Polytechnical University Xi'an, China(软件学院 西安理工大学 西安) Laboratoire L2Tl University Sorbonne Paris Nord Paris, France(L2Tl实验室 索邦巴黎北大学 巴黎)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15250 2025-09-23 cs.CV cs.AI 76%

Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token Pruning

Wenda Qin, Andrea Burns, Bryan A. Plummer, Margrit Betke

机构 * Boston University(波士顿大学)

专题命中 多模态训练与对齐 :multimodal(title);分类 cs.CV、cs.AI

Comments Accepted to EMNLP 2025. Data and code to be released at https://github.com/wdqin/VLN-NAP

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17747 2025-09-23 cs.CV cs.AI 73%

Dual-View Alignment Learning with Hierarchical-Prompt for Class-Imbalance Multi-Label Classification

Sheng Huang, Jiexuan Yan, Beiyan Liu, Bo Liu, Richang Hong

机构 * Ministry of Education Key Laboratory of Dependable Service Computing in Cyber Physical Society(教育部可信服务计算网络社会重点实验室) School of Big Data and Software Engineering(大数据与软件工程学院) School of Computer Science and Information Engineering(计算机科学与信息工程学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments accepted by IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16892 2025-09-23 cs.CV cs.AI 73%

Learning from Gene Names, Expression Values and Images: Contrastive Masked Text-Image Pretraining for Spatial Transcriptomics Representation Learning

Jiahe Qian, Yaoyu Fang, Ziqiao Weng, Xinkun Wang, Lee A. Cooper, Bo Zhou

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16957 2025-09-23 cs.CV 70%

MO R-CNN: Multispectral Oriented R-CNN for Object Detection in Remote Sensing Image

Leiyu Wang, Biao Jin, Feng Huang, Liqiong Chen, Zhengyong Wang, Xiaohai He, Honggang Chen

机构 * College of Electronics and Information Engineering, Sichuan University(四川大学电子信息工程学院) School of Mechanical Engineering and Automation, Fuzhou University(福州大学机械工程与自动化学院) Yunnan Key Laboratory of Software Engineering, Yunnan University(云南软件工程重点实验室)

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05255 2025-09-23 cs.CV cs.CL 62%

Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning

Yana Wei, Liang Zhao, Jianjian Sun, Kangheng Lin, Jisheng Yin, Jingcheng Hu, Yinmin Zhang, En Yu, Haoran Lv, Zejia Weng, Jia Wang, Chunrui Han, Yuang Peng, Qi Han, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Vishal M. Patel

机构 * Johns Hopkins University(约翰霍普金斯大学) StepFun BUPT(北京邮电大学) UCAS(中国科学院大学) THU(清华大学) HUST(华中科技大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17079 2025-09-23 cs.CV 57%

A Dual-Modulation Framework for RGB-T Crowd Counting via Spatially Modulated Attention and Adaptive Fusion

Yuhong Feng, Hongtao Chen, Qi Zhang, Jie Chen, Zhaoxi He, Mingzhe Liu, Jianghai Liao

机构 * College of Computer Science(计算机科学学院) Software Engineering, Shenzhen University, Shenzhen, China(软件工程,深圳大学,深圳,中国)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16348 2025-09-23 cs.AI 57%

A Unified AI Approach for Continuous Monitoring of Human Health and Diseases from Intensive Care Unit to Home with Physiological Foundation Models (UNIPHY+)

Minxiao Wang, Saurabh Kataria, Juntong Ni, Timothy G. Buchman, Jocelyn Grunwell, Mark Mai, Wei Jin, Matthew Clark, Stephanie Brown, Michael Fundora, Puneet Sharma, Tony Pan, Sam Khan, Timothy Ruchti, Naveen Muthu, Kevin Maher, Sivasubramanium V Bhavani, Xiao Hu

机构 * Emory University(埃默里大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏