arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-12 至 2025-08-12 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 16 篇

2508.07803 2025-08-12 cs.CV 83%

MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks

Yushen Xu, Xiaosong Li, Zhenyu Kuang, Xiaoqi Cheng, Haishu Tan, Huafeng Li

机构 * School of Physics and Optoelectronic Engineering(物理与光电工程学院) Guangdong-HongKong-Macao Joint Laboratory for Intelligent Micro-Nano Optoelectronic Technology(粤港澳联合智能微纳光电技术实验室) School of Information Engineering and Automation(信息工程与自动化学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07681 2025-08-12 cs.LG cs.AI 83%

MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation

Yooseok Lim, ByoungJun Jeon, Seong-A Park, Jisoo Lee, Sae Won Choi, Chang Wook Jeong, Ho-Geol Ryu, Hongyeol Lee, Hyun-Lim Yang

机构 * Seoul National University Hospital(首尔国立大学医院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06701 2025-08-12 cs.CV cs.AI cs.CL cs.LG cs.SD eess.AS 83%

MMFformer: Multimodal Fusion Transformer Network for Depression Detection

Md Rezwanul Haque, Md. Milon Islam, S M Taslim Uddin Raju, Hamdi Altaheri, Lobna Nassar, Fakhri Karray

机构 * Centre for Pattern Analysis and Machine Intelligence, Department of Electrical and Computer Engineering, University of Waterloo(模式分析与机器智能中心,电气与计算机工程系,滑铁卢大学) School of Engineering and Computing, Department of Computer Science and Engineering, American University of Ras Al Khaimah(工程与计算学院,计算机科学与工程系,阿联酋拉线哈姆斯美国大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted for the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00425 2025-08-12 cs.CV cs.AI 82%

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization

JiangYong Yu, Sifan Zhou, Dawei Yang, Shuo Wang, Shuoyu Li, Xing Hu, Chen Xu, Zukang Xu, Changyong Shu, Zhihang Yuan

机构 * Southeast University(东南大学) Xi'an Jiaotong University(西安交通大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM MM 2025. First PTQ solution for Multimodal large language models applicable to 5 mainstream MLLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06895 2025-08-12 cs.CV cs.AI 81%

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models

Jianting Tang, Yubo Wang, Haoyu Cao, Linli Xu

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00826 2025-08-12 cs.CL cs.AI cs.LG 81%

HERGC: Heterogeneous Experts Representation and Generative Completion for Multimodal Knowledge Graphs

Yongkang Xiao, Rui Zhang

机构 * University of Minnesota(明尼苏达大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04770 2025-08-12 cs.LG cs.AI q-bio.MN 79%

Bidirectional Hierarchical Protein Multi-Modal Representation Learning

Xuefeng Liu, Songhao Jiang, Chih-chan Tien, Jinbo Xu, Rick Stevens

机构 * Argonne National Laboratory(阿贡国家实验室)

专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06496 2025-08-12 cs.CV cs.MA 79%

Med-GRIM: Enhanced Zero-Shot Medical VQA using prompt-embedded Multimodal Graph RAG

Rakesh Raj Madavan, Akshat Kaimal, Hashim Faisal, Chandrakala S

机构 * Shiv Nadar University Chennai(施瓦斯纳大学钦奈)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07951 2025-08-12 cs.CV 79%

Scaling Laws for Native Multimodal Models

Mustafa Shukor, Enrico Fini, Victor Guilherme Turrisi da Costa, Matthieu Cord, Joshua Susskind, Alaaeldin El-Nouby

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2025 (Oral). 28 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07536 2025-08-12 cs.LG 78%

Physics-Informed Multimodal Bearing Fault Classification under Variable Operating Conditions using Transfer Learning

Tasfiq E. Alam, Md Manjurul Ahsan, Shivakumar Raman

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06566 2025-08-12 cs.CV cs.AI 73%

Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features

Manish Kansana, Elias Hossain, Shahram Rahimi, Noorbakhsh Amiri Golilarz

机构 * Department of Computer Science and Engineering, Mississippi State University(计算机科学与工程系,密苏里州立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06826 2025-08-12 cs.HC 67%

AdjustAR: AI-Driven In-Situ Adjustment of Site-Specific Augmented Reality Content

Nels Numan, Jessica Van Brummelen, Ziwen Lu, Anthony Steed

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract)

Comments 4 pages, 1 figure, ACM UIST 2025 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07804 2025-08-12 cs.CV 57%

Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning

Bao Li, Xiaomei Zhang, Miao Xu, Zhaoxin Fan, Xiangyu Zhu, Zhen Lei

机构 * CASIA(中国科学院自动化研究所) UCAS(中国科学院大学) CAIR, HKISI, CAS(中国科学院自动化研究所) Beihang University(北京航空航天大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07480 2025-08-12 eess.SP cs.AI cs.LG 57%

EEG-Language Pretraining for Highly Label-Efficient Clinical Phenotyping

Sam Gijsen, Kerstin Ritter

机构 * Charité – Universitätsmedizin Berlin, Department of Psychiatry and Psychotherapy, Berlin, Germany(柏林查理医院医学大学精神病与心理治疗系) Hertie Institute for AI in Brain Health, University of Tübingen, Germany(图宾根大学健康人工智能研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07738 2025-08-12 cs.LG 50%

Separation and Collaboration: Two-Level Routing Grouped Mixture-of-Experts for Multi-Domain Continual Learning

Jialu Zhou, Dianxi Shi, Shaowu Yang, Xinyu Wei, Mingyue Yang, Leqian Li, Mengzhu Wang, Chunping Qiu

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07717 2025-08-12 eess.SP 50%

Touch-Augmented Gaussian Splatting for Enhanced 3D Scene Reconstruction

Yuchen Gao, Xiao Xu, Eckehard Steinbach, Daniel E. Lucani, Qi Zhang

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏