arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6929 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6929 篇

2511.08152 2025-11-12 cs.CV cs.LG 79%

Boomda: Balanced Multi-objective Optimization for Multimodal Domain Adaptation

Jun Sun, Xinxin Zhang, Simin Hong, Jian Zhu, Xiang Gao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06805 2025-11-11 cs.AI cs.LG 79%

MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning

Jinhao Chen, Zhen Yang, Jianxin Shi, Tianyu Wo, Jie Tang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 19 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06593 2025-11-11 cs.CV 79%

Spatial-Frequency Enhanced Mamba for Multi-Modal Image Fusion

Hui Sun, Long Lv, Pingping Zhang, Tongdan Tang, Feng Tian, Weibing Sun, Huchuan Lu

机构 * School of Future Technology, School of Artificial Intelligence, Dalian University of Technology(大连理工大学未来技术学院、人工智能学院) Affiliated Zhongshan Hospital of Dalian University(大连大学附属中山医院) Central Hospital of Dalian University of Technology(大连理工大学中心医院) School of Information and Communication Engineering, Dalian University of Technology(大连理工大学信息与通信工程学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments This work is accepted by IEEE Transactions on Image Processing. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05474 2025-11-10 cs.CV 79%

Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection

Xian-Hong Huang, Hui-Kai Su, Chi-Chia Sun, Jun-Wei Hsieh

机构 * Department of Electrical Engineering, National Formosa University, Taiwan(台湾国立Formosa大学电子工程系) Department of Electrical Engineering, National Taipei University, Taiwan(台湾国立台北大学电子工程系) College of Artificial Intelligence and Green Energy, National Yang Ming Chiao Tung University, Taiwan(台湾国立阳明交通大学人工智能与再生能源学院)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04789 2025-11-06 cs.CV 79%

Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations

Gaia Di Lorenzo, Federico Tombari, Marc Pollefeys, Daniel Barath

机构 * ETH Zurich(苏黎世联邦理工学院) Google(谷歌) Microsoft(微软)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18174 2025-11-05 eess.SP cs.AI cs.LG 79%

NMCSE: Noise-Robust Multi-Modal Coupling Signal Estimation Method via Optimal Transport for Cardiovascular Disease Detection

Peihong Zhang, Zhixin Li, Rui Sang, Yuxuan Liu, Yiqiang Cai, Yizhou Tan, Shengchen Li

机构 * School of Advanced Technology, Xi’an Jiaotong-Liverpool University(先进技术学院,西安交通大学利物浦大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15991 2025-11-05 cs.CV 79%

CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D Detection

Huiming Yang, Wenzhuo Liu, Yicheng Qiao, Lei Yang, Xianzhu Zeng, Li Wang, Zhiwei Li, Zijian Zeng, Zhiying Jiang, Huaping Liu, Kunfeng Wang

机构 * Beijing University of Chemical Technology(北京化工大学) Division of Energy-Mobility Convergence, Beijing Institute of Technology(北京理工大学能源-交通融合学院) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动学院) School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与航空航天工程学院) School of Mechanical Engineering, Beijing Institute of Technology(北京理工大学机械工程学院) Institute of Computer Science and Digital Innovation, UCSI University(UCSI大学计算机科学与数字创新学院) State Key Laboratory of Intelligent Technology and Systems and Department of Computer Science and Technology, Tsinghua University(清华大学智能技术与系统国家重点实验室与计算机科学与技术系)

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01435 2025-11-04 cs.CV 79%

Contrast-Guided Cross-Modal Distillation for Thermal Object Detection

SiWoo Kim, JhongHyun An

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00859 2025-11-04 cs.CV 79%

Layer-Wise Modality Decomposition for Interpretable Multimodal Sensor Fusion

Jaehyun Park, Konyul Park, Daehun Kim, Junseo Park, Jun Won Choi

机构 * Seoul National University(首尔国立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19769 2025-11-04 cs.CV 79%

AIM: Adaptive Intra-Network Modulation for Balanced Multimodal Learning

Shu Shen, C. L. Philip Chen, Tong Zhang

机构 * Guangdong Provincial Key Laboratory of Computational AI Models and Cognitive Intelligence(广东省计算人工智能模型与认知智能重点实验室) School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院) Pazhou Lab(琶洲实验室) Engineering Research Center of the Ministry of Education on Health Intelligent Perception and Paralleled Digital-Human(教育部健康智能感知与并行数字人工程研究中心)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 13pages,7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27166 2025-11-03 cs.CV 79%

M^3Detection: Multi-Frame Multi-Level Feature Fusion for Multi-Modal 3D Object Detection with Camera and 4D Imaging Radar

Xiaozhi Li, Huijun Di, Jian Li, Feng Liu, Wei Liang

机构 * Radar Technology Research Institute, School of Information and Electronics, Beijing Institute of Technology(雷达技术研究所,信息电子学院,北京理工大学) Key Laboratory of Electronic and Information Technology in Satellite Navigation, Ministry of Education(卫星导航电子信息技术重点实验室,教育部) School of Computer Science, Beijing Institute of Technology(计算机学院,北京理工大学) Beijing Racobit Electronic Information Technology Co., Ltd.(北京Racobit电子信息技术有限公司)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24919 2025-10-30 cs.CV cs.LG 79%

Modality-Aware SAM: Sharpness-Aware-Minimization Driven Gradient Modulation for Harmonized Multimodal Learning

Hossein R. Nowdeh, Jie Ji, Xiaolong Ma, Fatemeh Afghah

机构 * Holcombe Department of ECE(霍尔科姆电气与计算机工程系) Clemson University(克莱姆森大学) University of Arizona(亚利桑那大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03318 2025-10-30 cs.CV 79%

Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning

Yibin Wang, Zhimin Li, Yuhang Zang, Chunyu Wang, Qinglin Lu, Cheng Jin, Jiaqi Wang

机构 * College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) Shanghai Innovation Institute(上海创新研究院) Shanghai AI Lab(上海人工智能实验室) Hunyuan, Tencent(腾讯 Hunyuan)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments [NeurIPS2025] Project Page: https://codegoat24.github.io/UnifiedReward/think

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06456 2025-10-29 cs.CV 79%

DynCIM: Dynamic Curriculum for Imbalanced Multimodal Learning

Chengxuan Qian, Kai Han, Jiaxin Liu, Zhenlong Yuan, Zhengzhong Zhu, Jingchao Wang, Chongwen Lyu, Jun Chen, Zhe Liu

机构 * Jiangsu University(江苏大学) UIUC(伊利诺伊大学香槟分校) Alibaba(阿里巴巴) UCAS(中国科学院大学) Sichuan University(四川大学) Peking University(北京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22964 2025-10-28 cs.CV 79%

Survey of Multimodal Geospatial Foundation Models: Techniques, Applications, and Challenges

Liling Yang, Ning Chen, Jun Yue, Yidan Liu, Jiayi Ma, Pedram Ghamisi, Antonio Plaza, Leyuan Fang

机构 * School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) Institute of Remote Sensing and Geographic Information System, Peking University(北京大学遥感与地理信息系统研究所) School of Automation, Central South University(中南大学自动化学院) Electronic Information School, Wuhan University(武汉大学电子信息学院) Helmholtz-Zentrum Dresden-Rossendorf(德累斯顿-罗斯托克亥姆霍尔茨中心) Lancaster Environment Centre, Lancaster University(兰卡斯特大学环境研究中心) Hyperspectral Computing Laboratory, Department of Technology of Computers and Communications, Escuela Politécnica, University of Extremadura(埃斯特雷马杜拉大学技术计算机与通讯系超光谱计算实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11128 2025-10-27 cs.LG cs.CV 79%

Lightweight Facial Landmark Detection in Thermal Images via Multi-Level Cross-Modal Knowledge Transfer

Qiyi Tong, Olivia Nocentini, Marta Lagomarsino, Kuanqi Cai, Marta Lorenzini, Arash Ajoudani

机构 * Human-Robot Interfaces and Interaction Laboratory, Istituto Italiano di Tecnologia, Genoa, Italy(人机交互实验室,意大利技术研究院,热那亚,意大利) Ph.D. Program of National Interest in Robotics and Intelligent Machines (DRIM), Università di Genova(机器人与智能机器国家利益博士项目,热那亚大学)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19201 2025-10-24 cs.CL 79%

DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding

Yunhai Hu, Tianhua Xia, Zining Liu, Rahul Raman, Xingyu Liu, Bo Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang

机构 * Courant Institute of Mathematical Sciences, New York University(纽约大学数学科学学院) Tandon School of Engineering, New York University(纽约大学工程学院) Cerebras Systems Inc.(Cerebras Systems公司) University of Pennsylvania(宾夕法尼亚大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18321 2025-10-22 cs.CV 79%

Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding

Jinlin Li, Yuran Wang, Yifei Yuan, Xiao Zhou, Yingying Zhang, Xixian Yong, Yefeng Zheng, Xian Wu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学 Gallup 学院) Department of Electrical and Computer Engineering, McGill University(麦吉尔大学电气与计算机工程系) School of Statistics, Renmin University of China(中国人民大学统计学院) Tencent Jarvis Lab(腾讯 Jarvis 实验室) Medical Artificial Intelligence Lab, Westlake University(西湖大学医学人工智能实验室)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18244 2025-10-22 cs.CV 79%

BlendCLIP: Bridging Synthetic and Real Domains for Zero-Shot 3D Object Classification with Multimodal Pretraining

Ajinkya Khoche, Gergő László Nagy, Maciej Wozniak, Thomas Gustafsson, Patric Jensfelt

机构 * KTH Royal Institute of Technology(皇家理工学院) Scania CV AB(斯堪尼亚公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12445 2025-10-22 cs.MM 79%

M3ST-DTI: A multi-task learning model for drug-target interactions based on multi-modal features and multi-stage alignment

Xiangyu Li, Ran Su, Liangliang Liu

专题命中 多模态训练与对齐 :multi-modal(title);cross-modal(abstract);分类 cs.MM

Comments This paper accepted by IEEE BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05914 2025-10-22 cs.CV 79%

Learning Collaborative Knowledge with Multimodal Representation for Polyp Re-Identification

Suncheng Xiang, Jiale Guan, Shilun Cai, Jiacheng Ruan, Dahong Qian

机构 * Shanghai Jiao Tong University(上海交通大学) Zhongshan Hospital of Fudan University(复旦大学中山医院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17394 2025-10-21 cs.LG cs.CV 79%

MILES: Modality-Informed Learning Rate Scheduler for Balancing Multimodal Learning

Alejandro Guerra-Manzanares, Farah E. Shamout

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted and presented at the 2025 International Joint Conference on Neural Networks (IJCNN'25). The paper was awarded an honorable mention (best 4 papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17289 2025-10-21 cs.CL 79%

Addressing Antisocial Behavior in Multi-Party Dialogs Through Multimodal Representation Learning

Hajar Bakarou, Mohamed Sinane El Messoussi, Anaïs Ollagnier

机构 * Universit\'e C \ te d'Azur, CNRS, Inria, I3S Sophia Antipolis France Universit\'e C \ te d'Azur, CNRS, Inria, I3S

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17078 2025-10-21 cs.CV 79%

Towards a Generalizable Fusion Architecture for Multimodal Object Detection

Jad Berjawi, Yoann Dupas, Christophe C'erin

机构 * Université Grenoble Alpes(格勒诺布尔大学) Université Sorbonne Paris Nord(巴黎-萨克勒大学) INRIA(法国国家信息与自动化研究所)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments 8 pages, 8 figures, accepted at ICCV 2025 MIRA Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08217 2025-10-21 cs.AI 79%

Quantum Federated Learning for Multimodal Data: A Modality-Agnostic Approach

Atit Pokharel, Ratun Rahman, Thomas Morris, Dinh C. Nguyen

机构 * Department of Electrical and Computer Engineering, The University of Alabama in Huntsville(电气与计算机工程系,阿拉巴马大学亨茨维尔分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments This paper was presented at BEAM with CVPR 2025

Journal ref Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 545-554. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15944 2025-10-21 cs.LG cs.AI 79%

Lyapunov-Stable Adaptive Control for Multimodal Concept Drift

Tianyu Bell Pan, Mengdi Zhu, Alexa Jordyn Cole, Ronald Wilson, Damon L. Woodard

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Florida Institute of National Security(佛罗里达国家安全研究所) Applied Artificial Intelligence Group(应用人工智能组) University of Florida(佛罗里达大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21514 2025-10-21 cs.CV 79%

G$^{2}$D: Boosting Multimodal Learning with Gradient-Guided Distillation

Mohammed Rakib, Arunkumar Bagavathi

机构 * Oklahoma State University(俄克拉荷马州立大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15317 2025-10-20 cs.AI 79%

VERITAS: Leveraging Vision Priors and Expert Fusion to Improve Multimodal Data

Tingqiao Xu, Ziru Zeng, Jiayu Chen

机构 * Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15026 2025-10-20 cs.CV 79%

MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning

Mattia Segu, Marta Tintore Gazulla, Yongqin Xian, Luc Van Gool, Federico Tombari

机构 * Google(谷歌) ETH Zurich(苏黎世联邦理工学院) INSAIT, Sofia University, St. Kliment Ohridski(INSAIT,索菲亚大学,圣克莱孟·奥赫里茨基)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14344 2025-10-17 cs.CR cs.AI 79%

BinCtx: Multi-Modal Representation Learning for Robust Android App Behavior Detection

Zichen Liu, Shao Yang, Xusheng Xiao

机构 * Arizona State University(亚利桑那州立大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏