arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2505.15804 2025-07-11 cs.CV 79%

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

Zongzhao Li, Zongyang Ma, Mingze Li, Songyou Li, Yu Rong, Tingyang Xu, Ziqi Zhang, Deli Zhao, Wenbing Huang

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) DAMO Academy, Alibaba Group, Hangzhou, China(阿里云达摩院) Hupan Lab, Hangzhou, China(杭州华普实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15285 2025-07-11 cs.CV 79%

EEPNet-V2: Patch-to-Pixel Solution for Efficient Cross-Modal Registration between LiDAR Point Cloud and Camera Image

Yuanchao Yue, Hui Yuan, Zhengxin Li, Shuai Li, Wei Zhang

机构 * School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17968 2025-07-10 cs.CV 79%

A Multimodal Fusion Framework for Bridge Defect Detection with Cross-Verification

Ravi Datta Rachuri, Duoduo Liao, Samhita Sarikonda, Datha Vaishnavi Kondur

机构 * School of Computing George Mason University Fairfax, VA(计算学院乔治·马歇尔大学弗劳伊德堡)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Big Data 2024

Journal ref 2024 IEEE International Conference on Big Data (BigData)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04945 2025-07-09 cs.CL cs.LG eess.SP 79%

MEIT: Multimodal Electrocardiogram Instruction Tuning on Large Language Models for Report Generation

Zhongwei Wan, Che Liu, Xin Wang, Chaofan Tao, Hui Shen, Jing Xiong, Rossella Arcucci, Huaxiu Yao, Mi Zhang

机构 * The Ohio State University(俄亥俄州立大学) Imperial College London(伦敦帝国理工学院) The University of Hong Kong(香港大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05165 2025-07-08 cs.CV 79%

Differential Attention for Multimodal Crisis Event Analysis

Nusrat Munia, Junfeng Zhu, Olfa Nasraoui, Abdullah-Al-Zubaer Imran

机构 * University of Kentucky(肯塔基大学) Kentucky Geological Survey(肯塔基地质调查局) University of Louisville(路易斯维尔大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Presented at CVPRw 2025, MMFM3

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04891 2025-07-08 eess.IV cs.CV 79%

MurreNet: Modeling Holistic Multimodal Interactions Between Histopathology and Genomic Profiles for Survival Prediction

Mingxin Liu, Chengfei Cai, Jun Li, Pengbo Xu, Jinze Li, Jiquan Ma, Jun Xu

机构 * Jiangsu Key Laboratory of Intelligent Medical Image Computing(江苏省智能医学图像计算重点实验室) School of Artificial Intelligence(人工智能学院) Nanjing University of Information Science and Technology(南京信息工程大学) College of Information Engineering(信息工程学院) Taizhou University(泰州大学) College of Bioinformatics Science and Technology(生物信息科学与技术学院) Harbin Medical University(哈尔滨医科大学) School of Computer and Big Data(计算机与大数据学院) Heilongjiang University(黑龙江大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 2 figures, Accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04369 2025-07-08 cs.CV 79%

MambaFusion: Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection

Hanshi Wang, Jin Gao, Weiming Hu, Zhipeng Zhang

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), CASIA(多模态人工智能系统国家重点实验室(MAIS),中国科学院自动化所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) Anyverse Intelligence Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information(北京多模态信息超级智能安全重点实验室) School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03917 2025-07-08 cs.LG cs.CV 79%

Consistency-Aware Padding for Incomplete Multi-Modal Alignment Clustering Based on Self-Repellent Greedy Anchor Search

Shubin Ma, Liang Zhao, Mingdong Lu, Yifan Guo, Bo Xu

机构 * School of Software Technology, Dalian University of Technology(大连理工大学软件学院)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments Accepted at IJCAI 2025. 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03304 2025-07-08 cs.CV 79%

Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations

Hai Huang, Yan Xia, Sashuai Zhou, Hanting Wang, Shulei Wang, Zhou Zhao

机构 * Zhejiang University(浙江大学)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02908 2025-07-08 cs.LG cs.AI 79%

Hyperbolic Kernel Graph Neural Networks for Neurocognitive Decline Analysis from Multimodal Brain Imaging

Meimei Yang, Yongheng Sun, Qianqian Wang, Andrea Bozoki, Maureen Kohi, Mingxia Liu

机构 * Department of Radiology and Biomedical Research Imaging Center (BRIC), University of North Carolina at Chapel Hill(放射科与生物医学研究成像中心(BRIC)、北卡罗来纳大学教堂山分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 14 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02205 2025-07-08 cs.CV 79%

Team RAS in 9th ABAW Competition: Multimodal Compound Expression Recognition Approach

Elena Ryumina, Maxim Markitantov, Alexandr Axyonov, Dmitry Ryumin, Mikhail Dolgushin, Alexey Karpov

机构 * St. Petersburg Federal Research Center of the Russian Academy of Sciences(俄罗斯科学院圣彼得堡联邦研究中心) ITMO University(ITMO大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 7

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05945 2025-07-04 cs.CV 79%

MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection

Zitian Wang, Zehao Huang, Yulu Gao, Naiyan Wang, Si Liu

机构 * Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02279 2025-07-04 cs.CV 79%

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

Juntao Liu, Liqiang Niu, Wenchao Chen, Jie Zhou, Fandong Meng

机构 * Pattern Recognition Center, WeChat AI, Tencent Inc(模式识别中心、微信AI、腾讯公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01984 2025-07-04 cs.LG cs.CL cs.SI 79%

Multimodal Misinformation Detection Using Early Fusion of Linguistic, Visual, and Social Features

Gautam Kishore Shahi

机构 * University of Duisburg-Essen(杜伊斯堡-埃森大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01054 2025-07-03 cs.LG cond-mat.mtrl-sci cs.AI 79%

XxaCT-NN: Structure Agnostic Multimodal Learning for Materials Science

Jithendaraa Subramanian, Linda Hung, Daniel Schweigert, Santosh Suram, Weike Ye

机构 * Toyota Research Institute(丰田研究机构)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04353 2025-07-03 cs.CV cs.LG 79%

DeFusion: An Effective Decoupling Fusion Network for Multi-Modal Pregnancy Prediction

Xueqiang Ouyang, Jia Wei, Wenjie Huo, Xiaocong Wang, Rui Li, Jianlong Zhou

机构 * School of Computer Science and Engineering, South China University of Technology(南方科技大学计算机科学与工程学院) Department of Obstetrics and Gynecology, Nanfang Hospital, Southern Medical University(南方医科大学 obstetrics and gynecology 部门) Golisano College of Computing and Information Sciences, Rochester Institute of Technology(罗切斯特理工学院 computing and information sciences 学院) UTS Data Science Institute, University of Technology Sydney(悉尼大学技术科学研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00849 2025-07-02 cs.CV 79%

UAVD-Mamba: Deformable Token Fusion Vision Mamba for Multimodal UAV Detection

Wei Li, Jiaman Tang, Yang Li, Beihao Xia, Ligang Tan, Hongmao Qin

机构 * College of Mechanical and Vehicle Engineering, Hunan University(湖南大学机械与车辆工程学院) School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments The paper was accepted by the 36th IEEE Intelligent Vehicles Symposium (IEEE IV 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00506 2025-07-02 cs.CV 79%

SCING:Towards More Efficient and Robust Person Re-Identification through Selective Cross-modal Prompt Tuning

Yunfei Xie, Yuxuan Cheng, Juncheng Wu, Haoyu Zhang, Yuyin Zhou, Shoudong Han

机构 * Huazhong University of Science and Technology(华中科技大学) Huazhong Agricultural University(华中农业大学) University of California, Santa Cruz(加州大学圣克ruz分校) City University of Hong Kong (Dongguan)(香港城市大学(东莞))

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16297 2025-07-01 cs.CV 79%

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

Renshan Zhang, Rui Shao, Gongwei Chen, Miao Zhang, Kaiwen Zhou, Weili Guan, Liqiang Nie

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to the IEEE/CVF International Conference on Computer Vision (ICCV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10557 2025-07-01 cs.CL 79%

MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models

Jianhong Tu, Zhuohao Ni, Nicholas Crispino, Zihao Yu, Michael Bendersky, Beliz Gunel, Ruoxi Jia, Xin Liu, Lingjuan Lyu, Dawn Song, Chenguang Wang

机构 * Washington University in St. Louis(华盛顿大学圣路易斯分校) The University of British Columbia(不列颠哥伦比亚大学) Google Research(谷歌研究) Virginia Tech(弗吉尼亚理工大学) University of California, Davis(加州大学戴维斯分校) Sony AI(索尼人工智能) University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16069 2025-06-30 cs.CV 79%

Disentangled and Interpretable Multimodal Attention Fusion for Cancer Survival Prediction

Aniek Eijpe, Soufyan Lakbir, Melis Erdal Cesur, Sara P. Oliveira, Sanne Abeln, Wilson Silva

机构 * AI Technology for Life, Department of Information and Computing Sciences, Department of Biology, Utrecht University, Utrecht, The Netherlands(AI技术与生命、信息与计算科学系、生物学系、乌得勒支大学、乌得勒支、荷兰) Computational Pathology group, Department of Pathology, The Netherlands Cancer Institute, Amsterdam, The Netherlands(计算病理组、病理学系、荷兰癌症研究所、阿姆斯特丹、荷兰)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 1 figure, 3 tables. Preprint submitted and accepted to MICCAI 2025. This preprint has not undergone peer review or any post-submission improvements or corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21393 2025-06-27 cs.AI 79%

TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding

Junwen Zhang, Pu Chen, Yin Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 43 pages and 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21237 2025-06-27 cs.CV 79%

DiMPLe -- Disentangled Multi-Modal Prompt Learning: Enhancing Out-Of-Distribution Alignment with Invariant and Spurious Feature Separation

Umaima Rahman, Mohammad Yaqub, Dwarikanath Mahapatra

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学) Khalifa University(卡比拉大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21018 2025-06-27 cs.CV 79%

LASFNet: A Lightweight Attention-Guided Self-Modulation Feature Fusion Network for Multimodal Object Detection

Lei Hao, Lina Xu, Chang Liu, Yanni Dong

机构 * School of Geophysics and Geomatics, China University of Geosciences, Wuhan(地质物理与地质信息学院,中国地质大学(武汉)) School of Resource and Environmental Sciences, Wuhan University(资源与环境科学学院,武汉大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19324 2025-06-25 cs.CV 79%

Memory-Augmented Incomplete Multimodal Survival Prediction via Cross-Slide and Gene-Attentive Hypergraph Learning

Mingcheng Qu, Guang Yang, Donglin Di, Yue Gao, Tonghua Su, Yang Song, Lei Fan

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) School of Software, Tsinghua University(清华大学软件学院) School of Computer Science and Engineering, UNSW Sydney(新南威尔士大学计算机科学与工程学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments accepted by MICCAI2025 code: https://github.com/MCPathology/M2Surv

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03507 2025-06-24 cs.CV 79%

Mineral segmentation using electron microscope images and spectral sampling through multimodal graph neural networks

Samuel Repka, Bořek Reich, Fedor Zolotarev, Tuomas Eerola, Pavel Zemčík

机构 * Lappeenranta-Lahti University of Technology(拉佩宁塔-拉赫蒂技术大学) Brno University of Technology - Faculty of Information Technology(布拉格技术大学-信息科技学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16879 2025-06-24 eess.IV cs.CV 79%

Ultra-high resolution multimodal MRI densely labelled holistic structural brain atlas

José V. Manjón, Sergio Morell-Ortega, Marina Ruiz-Perez, Boris Mansencal, Edern Le Bot, Marien Gadea, Enrique Lanuza, Gwenaelle Catheline, Thomas Tourdias, Vincent Planche, Rémi Giraud, Denis Rivière, Jean-François Mangin, Nicole Labra-Avila, Roberto Vivo-Hernando, Gregorio Rubio, Fernando Aparici, Maria de la Iglesia-Vaya, Pierrick Coupé

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17202 2025-06-23 cs.CV 79%

UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation

Teng Li, Quanfeng Lu, Lirui Zhao, Hao Li, Xizhou Zhu, Yu Qiao, Jun Zhang, Wenqi Shao

机构 * HKUST(香港科技大学) Shanghai AI Laboratory(上海人工智能实验室) SJTU(上海交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Code: https://github.com/tliby/UniFork

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17136 2025-06-23 cs.CV 79%

Semi-Supervised Multi-Modal Medical Image Segmentation for Complex Situations

Dongdong Meng, Sheng Li, Hao Wu, Guoping Wang, Xueqing Yan

机构 * School of Physics, Peking University, Beijing, China(北京大学物理系) School of Computer Science, Peking University, Beijing, China(北京大学计算机系) Department of Radiotherapy, Peking University Cancer Hospital, Beijing, China(北京大学肿瘤医院放疗科)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages, 2 figures, accepted at MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04443 2025-06-23 cs.AI 79%

POV Learning: Individual Alignment of Multimodal Models using Human Perception

Simon Werner, Katharina Christ, Laura Bernardy, Marion G. Müller, Achim Rettinger

机构 * Trier University(特里尔大学) University of Innsbruck(因斯布鲁克大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏