arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2504.16798 2025-04-24 cs.MM cs.CV cs.LG 81%

4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer's Diagnosis

Yuxiang Wei, Yanteng Zhang, Xi Xiao, Tianyang Wang, Xiao Wang, Vince D. Calhoun

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11477 2025-04-17 cs.CV cs.AI 81%

SDIGLM: Leveraging Large Language Models and Multi-Modal Chain of Thought for Structural Damage Identification

Yunkai Zhang, Shiyin Wei, Yong Huang, Yawu Su, Shanshan Lu, Hui Li

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11459 2025-04-17 cs.AI cs.CL cs.IR 81%

From Conceptual Data Models to Multimodal Representation

Peter Stockinger

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments in French language

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11082 2025-04-16 cs.CL cs.AI 81%

DeepMLF: Multimodal language model with learnable tokens for deep fusion in sentiment analysis

Efthymios Georgiou, Vassilis Katsouros, Yannis Avrithis, Alexandros Potamianos

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05158 2025-04-08 cs.SD cs.AI eess.AS 81%

Leveraging Label Potential for Enhanced Multimodal Emotion Recognition

Xuechun Shao, Yinfeng Yu, Liejun Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI、eess.AS

Comments Main paper (8 pages). Accepted for publication by IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03735 2025-04-08 cs.CR cs.AI cs.CL cs.CY cs.LG 81%

Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots

Erfan Shayegani, G M Shahariar, Sara Abdali, Lei Yu, Nael Abu-Ghazaleh, Yue Dong

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02175 2025-04-03 cs.CV cs.AI cs.LG 81%

DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models

Saeed Ranjbar Alvar, Gursimran Singh, Mohammad Akbari, Yong Zhang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15362 2025-03-26 cs.CV cs.AI 81%

A Multimodal Knowledge-enhanced Whole-slide Pathology Foundation Model

Yingxue Xu, Yihui Wang, Fengtao Zhou, Jiabo Ma, Cheng Jin, Shu Yang, Jinbang Li, Zhengyu Zhang, Chenglong Zhao, Huajun Zhou, Zhenhui Li, Huangjing Lin, Xin Wang, Jiguang Wang, Anjia Han, Ronald Cheong Kin Chan, Li Liang, Xiuming Zhang, Hao Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 62 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01653 2025-03-25 cs.LG cs.AI cs.CV 81%

Distilled Prompt Learning for Incomplete Multimodal Survival Prediction

Yingxue Xu, Fengtao Zhou, Chenyu Zhao, Yihui Wang, Can Yang, Hao Chen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17777 2025-03-24 cs.AI cs.CV cs.LG eess.SP 81%

Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality Alignment

Shenghong Dai, Shiqi Jiang, Yifan Yang, Ting Cao, Mo Li, Suman Banerjee, Lili Qiu

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by SenSys'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10371 2025-03-14 cs.CV cs.AI cs.LG 81%

A Multimodal Fusion Model Leveraging MLP Mixer and Handcrafted Features-based Deep Learning Networks for Facial Palsy Detection

Heng Yim Nicole Oo, Min Hun Lee, Jeong Hoon Lim

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments PAKDD 2025. arXiv admin note: text overlap with arXiv:2405.16496

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16496 2025-03-14 cs.CV cs.AI cs.LG 81%

Exploring a Multimodal Fusion-based Deep Learning Network for Detecting Facial Palsy

Heng Yim Nicole Oo, Min Hun Lee, Jeong Hoon Lim

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments IJCAI 2024 4th AI for Ageless Aging Workshop (AIAA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07943 2025-03-12 cs.CV cs.CL 81%

Enhancing Sentiment Analysis through Multimodal Fusion: A BERT-DINOv2 Approach

Taoxu Zhao, Meisi Li, Kehao Chen, Liye Wang, Xucheng Zhou, Kunal Chaturvedi, Mukesh Prasad, Ali Anaissi, Ali Braytee

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02849 2025-03-05 cs.CV cs.AI 81%

Multimodal Deep Learning for Subtype Classification in Breast Cancer Using Histopathological Images and Gene Expression Data

Amin Honarmandi Shandiz

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13085 2025-03-04 cs.LG cs.CL cs.CV 81%

MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models

Peng Xia, Kangyu Zhu, Haoran Li, Tianze Wang, Weijia Shi, Sheng Wang, Linjun Zhang, James Zou, Huaxiu Yao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01506 2025-03-03 cs.CV cs.AI cs.LG 81%

Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion

Dexuan Ding, Lei Wang, Liyun Zhu, Tom Gedeon, Piotr Koniusz

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at the Thirteenth International Conference on Learning Representations (ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09925 2025-02-17 cs.CV cs.AI 81%

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types

Jiankang Chen, Tianke Zhang, Changyi Liu, Haojie Ding, Yaya Shi, Feng Cheng, Huihui Xiao, Bin Wen, Fan Yang, Tingting Gao, Di Zhang

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09675 2025-02-17 cs.CL cs.AI cs.LG 81%

Multi-level Conflict-Aware Network for Multi-modal Sentiment Analysis

Yubo Gao, Haotian Wu, Lei Zhang

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CL、cs.AI

Comments 5 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04794 2025-02-17 eess.IV cs.AI cs.CV 81%

MedMimic: Physician-Inspired Multimodal Fusion for Early Diagnosis of Fever of Unknown Origin

Minrui Chen, Yi Zhou, Huidong Jiang, Yuhan Zhu, Guanjie Zou, Minqi Chen, Rong Tian, Hiroto Saigo

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11959 2025-02-13 cs.CV cs.AI cs.LG 81%

Gramian Multimodal Representation Learning and Alignment

Giordano Cicchetti, Eleonora Grassucci, Luigi Sigillo, Danilo Comminiello

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01158 2025-02-05 cs.LG cs.AI cs.CV 81%

MIND: Modality-Informed Knowledge Distillation Framework for Multimodal Clinical Prediction Tasks

Alejandro Guerra-Manzanares, Farah E. Shamout

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Published in Transactions on Machine Learning Research (01/2025), https://openreview.net/forum?id=BhOJreYmur&noteId=ymnAhncuez

Journal ref Transactions on Machine Learning Research (TMLR), 01/2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01535 2025-02-04 cs.CV cs.CL q-bio.QM 81%

VisTA: Vision-Text Alignment Model with Contrastive Learning using Multimodal Data for Evidence-Driven, Reliable, and Explainable Alzheimer's Disease Diagnosis

Duy-Cat Can, Linh D. Dang, Quang-Huy Tang, Dang Minh Ly, Huong Ha, Guillaume Blanc, Oliver Y. Chén, Binh T. Nguyen

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18124 2025-02-04 cs.CV cs.AI 81%

REMOTE: Real-time Ego-motion Tracking for Various Endoscopes via Multimodal Visual Feature Learning

Liangjing Shao, Benshuang Chen, Shuting Zhao, Xinrong Chen

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18670 2025-02-03 cs.CV cs.AI 81%

High-Accuracy ECG Image Interpretation using Parameter-Efficient LoRA Fine-Tuning with Multimodal LLaMA 3.2

Nandakishor M, Anjali M

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17699 2025-01-30 eess.IV cs.AI cs.CV 81%

PulmoFusion: Advancing Pulmonary Health with Efficient Multi-Modal Fusion

Ahmed Sharshar, Yasser Attia, Mohammad Yaqub, Mohsen Guizani

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Journal ref (ISBI 2025) 2025 IEEE International Symposium on Biomedical Imaging

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02301 2025-01-30 cs.AI cs.CL cs.CR 81%

Large Multimodal Agents for Accurate Phishing Detection with Enhanced Token Optimization and Cost Reduction

Fouad Trad, Ali Chehab

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted in the 2nd International Conference on Foundation and Large Language Models (FLLM2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05803 2025-01-28 cs.CV cs.AI 81%

Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Zhihang Lin, Mingbao Lin, Luxi Lin, Rongrong Ji

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13988 2025-01-27 cs.RO cs.AI cs.CV 81%

MCRL4OR: Multimodal Contrastive Representation Learning for Off-Road Environmental Perception

Yi Yang, Zhang Zhang, Liang Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Github repository: https://github.com/1uciusy/MCRL4OR

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13985 2025-01-27 cs.LG cs.AI cs.CV 81%

Pilot: Building the Federated Multimodal Instruction Tuning Framework

Baochen Xiong, Xiaoshan Yang, Yaguang Song, Yaowei Wang, Changsheng Xu

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08585 2025-01-16 eess.SP cs.AI cs.CV cs.LG 81%

A Systematic Review of Machine Learning Methods for Multimodal EEG Data in Clinical Application

Siqi Zhao, Wangyang Li, Xiru Wang, Stevie Foglia, Hongzhao Tan, Bohan Zhang, Ameer Hamoodi, Aimee Nelson, Zhen Gao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments This paper includes 4 figures, 6 tables, and totals 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏