arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2508.08279 2025-08-13 cs.LG cs.AI 79%

XFMNet: Decoding Cross-Site and Nonstationary Water Patterns via Stepwise Multimodal Fusion for Long-Term Water Quality Forecasting

Ziqi Wang, Hailiang Zhao, Cheng Bao, Wenzhuo Qian, Yuhao Yang, Xueqiang Sun, Shuiguang Deng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04770 2025-08-12 cs.LG cs.AI q-bio.MN 79%

Bidirectional Hierarchical Protein Multi-Modal Representation Learning

Xuefeng Liu, Songhao Jiang, Chih-chan Tien, Jinbo Xu, Rick Stevens

机构 * Argonne National Laboratory(阿贡国家实验室)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06496 2025-08-12 cs.CV cs.MA 79%

Med-GRIM: Enhanced Zero-Shot Medical VQA using prompt-embedded Multimodal Graph RAG

Rakesh Raj Madavan, Akshat Kaimal, Hashim Faisal, Chandrakala S

机构 * Shiv Nadar University Chennai(施瓦斯纳大学钦奈)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07951 2025-08-12 cs.CV 79%

Scaling Laws for Native Multimodal Models

Mustafa Shukor, Enrico Fini, Victor Guilherme Turrisi da Costa, Matthieu Cord, Joshua Susskind, Alaaeldin El-Nouby

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2025 (Oral). 28 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05934 2025-08-11 cs.HC cs.AI cs.LG 79%

ASLSL: Adaptive shared latent structure learning with incomplete multi-modal physiological data for multi-dimensional emotional feature selection

Xueyuan Xu, Tianze Yu, Wenjia Dong, Fulin Wei, Li Zhuo

机构 * School of Information Science and Technology, Beijing University of Technology, Beijing 100124, China(信息科学与技术学院,北京理工大学,北京) School of Artificial Intelligence, Anhui University, Beijing 100124, China(人工智能学院,安徽大学,北京)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03155 2025-08-11 cs.LG cs.AI 79%

Fusing Cross-Domain Knowledge from Multimodal Data to Solve Problems in the Physical World

Yu Zheng

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06836 2025-08-11 cs.LG cond-mat.mtrl-sci cs.AI 79%

CAST: Cross Attention based multimodal fusion of Structure and Text for materials property prediction

Jaewan Lee, Changyoung Park, Hongjun Yang, Sungbin Lim, Woohyung Lim, Sehui Han

机构 * LG AI Research(LG人工智能研究)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04205 2025-08-07 cs.CV 79%

Small Lesions-aware Bidirectional Multimodal Multiscale Fusion Network for Lung Disease Classification

Jianxun Yu, Ruiquan Ge, Zhipeng Wang, Cheng Yang, Chenyu Lin, Xianjun Fu, Jikui Liu, Ahmed Elazab, Changmiao Wang

机构 * Xidian University(西安电子科技大学) Hangzhou Dianzi University(杭州电子科技大学) Zhejiang College of Security Technology, School of Artificial Intelligence(浙江安全技术学院人工智能学院) Shenzhen Polytechnic University(深圳职业技术学院) Shenzhen University(深圳大学) Shenzhen Research Institute of Big Data(深圳大数据研究院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08344 2025-08-06 cs.CV 79%

MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion

Jihao Gu, Fei Wang, Kun Li, Yanyan Wei, Zhiliang Wu, Dan Guo

机构 * University College London (UCL)(伦敦大学学院) School of Computer Science(计算机科学学院) Information Engineering, School of Artificial Intelligence, Hefei University of Technology (HFUT)(信息工程学院,人工智能学院,合肥工业大学) ReLER, CCAI, Zhejiang University, China(ReLER、CCAI、浙江大学,中国) Key Laboratory of Knowledge Engineering with Big Data (HFUT), Ministry of Education(大数据知识工程重点实验室(HFUT),教育部) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China(人工智能研究所,合肥综合性国家科学中心,中国) Xinsight Lab, Research Institute, Hefei Zhongjuyuan Intelligent Technology Co., Ltd., China(Xinsight实验室,研究院,合肥中睿智能科技有限公司,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 1st Place in Micro-gesture Classification sub-challenge in 3rd MiGA at IJCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01316 2025-08-05 cs.CV cs.HC 79%

Multimodal Attention-Aware Fusion for Diagnosing Distal Myopathy: Evaluating Model Interpretability and Clinician Trust

Mohsen Abbaspour Onari, Lucie Charlotte Magister, Yaoxin Wu, Amalia Lupi, Dario Creazzo, Mattia Tordin, Luigi Di Donatantonio, Emilio Quaia, Chao Zhang, Isel Grau, Marco S. Nobile, Yingqian Zhang, Pietro Liò

机构 * Information Systems Group, Eindhoven University of Technology, The Netherlands(埃因霍温技术大学信息系统组) Eindhoven Artificial Intelligence Systems Institute, The Netherlands(埃因霍温人工智能系统研究所) Department of Computer Science and Technology, University of Cambridge, United Kingdom(剑桥大学计算机科学与技术系) Department of Medicine - DIMED, Padua University Hospital, Italy(帕多瓦大学医院医学部-DIMED) Department of Environmental Sciences, Informatics, and Statistics, Ca’ Foscari University of Venice, Italy(威尼斯卡弗里大学环境科学、信息学与统计学系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00963 2025-08-05 cs.LG cs.AI 79%

Rethinking Multimodality: Optimizing Multimodal Deep Learning for Biomedical Signal Classification

Timothy Oladunni, Alex Wong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00830 2025-08-05 cs.CV 79%

Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

Yang Cao, Yihan Zeng, Hang Xu, Dan Xu

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科技大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

Comments Code Page: NeurIPS2023" target="_blank" rel="noopener">https://github.com/yangcaoai/CoDA_NeurIPS2023 This paper is accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.07348 2025-08-05 cs.CV 79%

Juggling With Representations: On the Information Transfer Between Imagery, Point Clouds, and Meshes for Multi-Modal Semantics

Dominik Laupheimer, Norbert Haala

机构 * Institute for Photogrammetry, University of Stuttgart, Germany(摄影测量研究所,斯图加特大学,德国)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00447 2025-08-04 cs.CV cs.LG 79%

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text

Anju Rani, Daniel Ortiz-Arroyo, Petar Durdevic

机构 * Department of Energy Technology(能源技术系) Aalborg University(奥尔堡大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22268 2025-08-04 cs.IR cs.AI 79%

Multi-modal Relational Item Representation Learning for Inferring Substitutable and Complementary Items

Junting Wang, Chenghuan Guo, Jiao Yang, Yanhui Guo, Yan Gao, Hari Sundaram

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19651 2025-08-04 cs.LG cs.CL 79%

Unlocking Multi-Modal Potentials for Link Prediction on Dynamic Text-Attributed Graphs

Yuanyuan Xu, Wenjie Zhang, Ying Zhang, Xuemin Lin, Xiwei Xu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22037 2025-07-30 cs.CR cs.AI 79%

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security

Muzhi Dai, Shixuan Liu, Zhiyuan Zhao, Junyu Gao, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI),中国电信,中国) Northwestern Polytechnical University(西北工业大学) China Telecom, China(中国电信,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10774 2025-07-30 cs.LG cs.AI 79%

Context-Aware Probabilistic Modeling with LLM for Multimodal Time Series Forecasting

Yueyang Yao, Jiajun Li, Xingyuan Dai, MengMeng Zhang, Xiaoyan Gong, Fei-Yue Wang, Yisheng Lv

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20714 2025-07-29 cs.LG cs.AI q-bio.QM stat.AP 79%

Prostate Cancer Classification Using Multimodal Feature Fusion and Explainable AI

Asma Sadia Khan, Fariba Tasnia Khan, Tanjim Mahmud, Salman Karim Khan, Rishita Chakma, Nahed Sharmen, Mohammad Shahadat Hossain, Karl Andersson

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14683 2025-07-29 cs.CV 79%

Emerging Properties in Unified Multimodal Pretraining

Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, Guang Shi, Haoqi Fan

机构 * ByteDance Seed(字节跳动种子) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Monash University(墨尔本大学) Hong Kong University of Science and Technology(香港科学与技术大学) UC Santa Cruz(加州大学圣克ruz分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 37 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19253 2025-07-28 cs.CV 79%

BridgeNet: A Unified Multimodal Framework for Bridging 2D and 3D Industrial Anomaly Detection

An Xiang, Zixuan Huang, Xitong Gao, Kejiang Ye, Cheng-zhong Xu

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Shenzhen University of Advanced Technology(深圳先进技术大学) State Key Lab of IOTSC, Department of CIS, University of Macau(物联网科学国家重点实验室,澳门大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19052 2025-07-28 cs.CV 79%

Probing Multimodal Fusion in the Brain: The Dominance of Audiovisual Streams in Naturalistic Encoding

Hamid Abdollahi, Amir Hossein Mansouri Majoumerd, Amir Hossein Bagheri Baboukani, Amir Abolfazl Suratgar, Mohammad Bagher Menhaj

机构 * Distributed and Intelligent Optimization Research Laboratory(分布式智能优化研究实验室) Electrical Engineering Department(电气工程系) Amirkabir University of Technology(阿米尔卡比尔技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10203 2025-07-22 cs.CV 79%

Improving Multimodal Learning via Imbalanced Learning

Shicai Wei, Chunbo Luo, Yang Luo

机构 * University of Electronic Science and Technology of China(电子科技大学) National and Local Joint Engineering Research Center for Cloud Operating System(云计算操作系统国家级地方联合工程研究中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14175 2025-07-22 cs.LG cs.AI stat.AP 79%

Latent Space Data Fusion Outperforms Early Fusion in Multimodal Mental Health Digital Phenotyping Data

Youcef Barkat, Dylan Hamitouche, Deven Parekh, Ivy Guo, David Benrimoh

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13026 2025-07-17 cs.CV 79%

HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model

Tao Wang, Changxu Cheng, Lingfeng Wang, Senda Chen, Wuyue Zhao

机构 * Uni-Ubi Zhejiang University(浙江大学) Tongji University(同济大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ICCV 2025; the code is at https://github.com/yayafengzi/LMM-HiMTok

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11003 2025-07-16 cs.CV 79%

Bridge Feature Matching and Cross-Modal Alignment with Mutual-filtering for Zero-shot Anomaly Detection

Yuhu Bai, Jiangning Zhang, Yunkang Cao, Guangyuan Lu, Qingdong He, Xiangtai Li, Guanzhong Tian

机构 * Zhejiang University(浙江大学) YouTu Lab, Tencent(腾讯YouTu实验室) Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10620 2025-07-16 cs.LG cs.AI 79%

LLMs Meet Cross-Modal Time Series Analytics: Overview and Directions

Chenxi Liu, Hao Miao, Cheng Long, Yan Zhao, Ziyue Li, Panos Kalnis

机构 * Nanyang Technological University(南洋理工大学) The Hong Kong Polytechnic University(香港理工大学) University of Electronic Science and Technology of China(电子科学与技术大学) Technical University of Munich(慕尼黑技术大学) King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.AI

Comments Accepted at SSTD 2025 (Tutorial). arXiv admin note: text overlap with arXiv:2505.02583

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09334 2025-07-15 cs.CV 79%

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding

Wencan Huang, Daizong Liu, Wei Hu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10715 2025-07-14 cs.CV 79%

EVT: Efficient View Transformation for Multi-Modal 3D Object Detection

Yongjin Lee, Hyeon-Mun Jeong, Yurim Jeon, Sanghyun Kim

机构 * ThorDrive Co., Ltd(ThorDrive公司) Seoul National University(首尔国立大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08000 2025-07-11 cs.CV cs.LG 79%

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models

Helen Qu, Sang Michael Xie

机构 * Flatiron Institute(Flatiron研究所) Stanford University(斯坦福大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏