arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-28 至 2025-10-28 共收录 15 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 15 篇

2501.17823 2025-10-28 cs.CV cs.AI cs.LG 88%

Robust Multimodal Learning via Cross-Modal Proxy Tokens

Md Kaykobad Reza, Ameya Patil, Mashhour Solh, M. Salman Asif

机构 * University of California Riverside(加州大学河滨分校) Amazon(亚马逊)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 28 Pages, 13 Figures, 11 Tables. Accepted by Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22829 2025-10-28 cs.CV cs.AI cs.MM 85%

LLM-based Fusion of Multi-modal Features for Commercial Memorability Prediction

Aleksandar Pramov

机构 * Georgia Institute of Technology, USA(佐治亚理工学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21793 2025-10-28 cs.CV cs.AI eess.IV 84%

2D_3D Feature Fusion via Cross-Modal Latent Synthesis and Attention Guided Restoration for Industrial Anomaly Detection

Usman Ali, Ali Zia, Abdul Rehman, Umer Ramzan, Zohaib Hassan, Talha Sattar, Jing Wang, Wei Xiang

机构 * GIFT University(GIFT大学) La Trobe University(拉特罗布大学) Department of Primary Industries(初级产业部门)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at 26th International Conference on Digital Image Computing: Techniques and Applications (DICTA 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23273 2025-10-28 cs.LG cs.AI q-bio.QM 83%

A Novel Framework for Multi-Modal Protein Representation Learning

Runjie Zheng, Zhen Wang, Anjie Qiao, Jiancong Xie, Jiahua Rao, Yuedong Yang

机构 * School of Computer Science and Engineering, Sun Yat-sen University (SYSU)(计算机科学与工程学院,中山大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 35 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23151 2025-10-28 cs.CV cs.LG 83%

AG-Fusion: adaptive gated multimodal fusion for 3d object detection in complex scenes

Sixian Liu, Chen Xu, Qiang Wang, Donghai Shi, Yiwen Li

机构 * Yaowu Technology Co., Ltd, Shenzhen, China(深圳优华科技有限公司,深圳,中国)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21829 2025-10-28 cs.CV 83%

A Flow Model with Low-Rank Transformers for Incomplete Multimodal Survival Analysis

Yi Yin, Yuntao Shou, Zao Dai, Yun Peng, Tao Meng, Wei Ai, Keqin Li

机构 * College of Computer and Mathematics, Central South University of Forestry and Technology(计算机与数学学院,中央南大学林业与技术大学) School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) Department of Computer Science, State University of New York(计算机科学系,纽约州立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22507 2025-10-28 cs.CV cs.AI 81%

GateFuseNet: An Adaptive 3D Multimodal Neuroimaging Fusion Network for Parkinson's Disease Diagnosis

Rui Jin, Chen Chen, Yin Liu, Hongfu Sun, Min Zeng, Min Li, Yang Gao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments The first two authors contributed equally to this work. Correspondence to: Yang Gao, E-mail: yang.gao@csu.edu.cn

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22964 2025-10-28 cs.CV 79%

Survey of Multimodal Geospatial Foundation Models: Techniques, Applications, and Challenges

Liling Yang, Ning Chen, Jun Yue, Yidan Liu, Jiayi Ma, Pedram Ghamisi, Antonio Plaza, Leyuan Fang

机构 * School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) Institute of Remote Sensing and Geographic Information System, Peking University(北京大学遥感与地理信息系统研究所) School of Automation, Central South University(中南大学自动化学院) Electronic Information School, Wuhan University(武汉大学电子信息学院) Helmholtz-Zentrum Dresden-Rossendorf(德累斯顿-罗斯托克亥姆霍尔茨中心) Lancaster Environment Centre, Lancaster University(兰卡斯特大学环境研究中心) Hyperspectral Computing Laboratory, Department of Technology of Computers and Communications, Escuela Politécnica, University of Extremadura(埃斯特雷马杜拉大学技术计算机与通讯系超光谱计算实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21808 2025-10-28 cs.CV cs.AI 62%

Semantic Relation-Enhanced CLIP Adapter for Domain Adaptive Zero-Shot Learning

Jiaao Yu, Mingjie Han, Jinkun Jiang, Junyu Dong, Tao Gong, Man Lan

机构 * School of Computer Science and Technology, East China Normal University, China(东华大学计算机科学与技术学院) College of Computer Science and Technology, Ocean University of China, China(中国海洋大学计算机科学与技术学院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21794 2025-10-28 cs.CV cs.AI 62%

Token-Level Inference-Time Alignment for Vision-Language Models

Kejia Chen, Jiawen Zhang, Jiacong Hu, Kewei Gao, Jian Lou, Zunlei Feng, Mingli Song

机构 * Zhejiang University(浙江大学) Sun Yat-sen University(中山大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03201 2025-10-28 cs.CV 57%

AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding

Yidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao, Nicu Sebe

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Trento(特伦托大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22301 2025-10-28 cs.LG cs.AI 57%

AnyECG-Lab: An Exploration Study of Fine-tuning an ECG Foundation Model to Estimate Laboratory Values from Single-Lead ECG Signals

Yujie Xiao, Gongzhen Tang, Wenhui Liu, Jun Li, Guangkun Nie, Zhuoran Kan, Deyun Zhang, Qinghao Zhao, Shenda Hong

机构 * Institute of Medical Technology, Peking University Health Science Center(北京大学人民医院医学技术研究所) National Institute of Health Data Science, Peking University(北京大学国家健康数据科学研究院) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) HeartVoice Medical Technology(心声医疗技术) Department of Cardiology, Peking University People’s Hospital(北京大学人民医院心内科) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) State Key Laboratory of Vascular Homeostasis and Remodeling, NHC Key Laboratory of Cardiovascular Molecular Biology and Regulatory Peptides, Peking University(国家心血管病分子生物学与调节肽重点实验室,北京大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16188 2025-10-28 cs.CV 57%

Think or Not Think: A Study of Explicit Thinking in Rule-Based Visual Reinforcement Fine-Tuning

Ming Li, Jike Zhong, Shitian Zhao, Yuxiang Lai, Haoquan Zhang, Wang Bill Zhu, Kaipeng Zhang

机构 * Shanghai AI Laboratory(上海人工智能实验室) University of Southern California(南加州大学) Emory University(埃默里大学) Chinese University of Hong Kong(香港中文大学)

专题命中 多模态训练与对齐 :MLLM(abstract);分类 cs.CV

Comments Neurips 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23449 2025-10-28 cs.LG 50%

Schrodinger Neural Network and Uncertainty Quantification: Quantum Machine

M. M. Hammad

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments 29 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22239 2025-10-28 eess.IV cs.LG q-bio.QM 50%

Synthetic-to-Real Transfer Learning for Chromatin-Sensitive PWS Microscopy

Jahidul Arafat, Sanjaya Poudel

机构 * Department of Computer Science and Software Engineering, Auburn University, Alabama, USA(计算机科学与软件工程系,阿伯茨罕大学,阿拉巴马州,美国)

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments 24 pages, 5 figures and 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏