arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-11 至 2025-11-11 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 16 篇

2503.18135 2025-11-11 cs.CV 88%

MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation

Jiaxin Huang, Runnan Chen, Ziwen Li, Zhengqing Gao, Xiao He, Yandong Guo, Mingming Gong, Tongliang Liu

机构 * MBZUAI The University of Sydney(悉尼大学) The University of Melbourne(墨尔本大学) AI2Robotic

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CV

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07274 2025-11-11 cs.LG 82%

Multi-modal Dynamic Proxy Learning for Personalized Multiple Clustering

Jinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li, Ziyue Peng, Zewei Liu, Hewei Wang, Jiayi Zhang, Edith C. H. Ngai

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract)

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06805 2025-11-11 cs.AI cs.LG 79%

MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning

Jinhao Chen, Zhen Yang, Jianxin Shi, Tianyu Wo, Jie Tang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 19 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06593 2025-11-11 cs.CV 79%

Spatial-Frequency Enhanced Mamba for Multi-Modal Image Fusion

Hui Sun, Long Lv, Pingping Zhang, Tongdan Tang, Feng Tian, Weibing Sun, Huchuan Lu

机构 * School of Future Technology, School of Artificial Intelligence, Dalian University of Technology(大连理工大学未来技术学院、人工智能学院) Affiliated Zhongshan Hospital of Dalian University(大连大学附属中山医院) Central Hospital of Dalian University of Technology(大连理工大学中心医院) School of Information and Communication Engineering, Dalian University of Technology(大连理工大学信息与通信工程学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments This work is accepted by IEEE Transactions on Image Processing. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06723 2025-11-11 cs.LG 78%

Multi-Modal Continual Learning via Cross-Modality Adapters and Representation Alignment with Knowledge Preservation

Evelyn Chee, Wynne Hsu, Mong Li Lee

机构 * School of Computing, National University of Singapore(computing学院,新加坡国立大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments Accepted to ECAI 2025

Journal ref 28th European Conference on Artificial Intelligence (ECAI), 2025, pp.1083-1090

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05716 2025-11-11 cs.LG 78%

Distributionally Robust Multimodal Machine Learning

Peilin Yang, Yu Ma

机构 * University of Wisconsin, Madison(威斯康星大学麦迪逊分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05553 2025-11-11 cs.CV cs.AI 73%

EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning

Xinyan Cai, Shiguang Wu, Dafeng Chi, Yuzheng Zhuang, Xingyue Quan, Jianye Hao, Qiang Guan

机构 * Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00801 2025-11-11 cs.CV cs.AI cs.RO 73%

Environment-Driven Online LiDAR-Camera Extrinsic Calibration

Zhiwei Huang, Jiaqi Li, Hongbo Zhao, Xiao Ma, Ping Zhong, Xiaohu Zhou, Wei Ye, Rui Fan

机构 * Department of Control Science & Engineering, the College of Electronic & Information Engineering, Tongji University(控制科学与工程系,电子与信息工程学院,同济大学) School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学) Beijing Institute of Aerospace Control Devices(北京航天控制器件研究所) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06225 2025-11-11 cs.CV 70%

MoRA: Missing Modality Low-Rank Adaptation for Visual Recognition

Shu Zhao, Nilesh Ahuja, Tan Yu, Tianyi Shen, Vijaykrishnan Narayanan

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Intel(英特尔) NVIDIA(英伟达)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06836 2025-11-11 cs.CV cs.AI 62%

NeuroBridge: Bio-Inspired Self-Supervised EEG-to-Image Decoding via Cognitive Priors and Bidirectional Semantic Alignment

Wenjiang Zhang, Sifeng Wang, Yuwei Su, Xinyu Li, Chen Zhang, Suyu Zhong

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19856 2025-11-11 cs.CV cs.AI 62%

RaGS: Unleashing 3D Gaussian Splatting from 4D Radar and Monocular Cues for 3D Object Detection

Xiaokai Bai, Chenxu Zhou, Lianqing Zheng, Si-Yuan Cao, Jianan Liu, Xiaohan Zhang, Yiming Li, Zhengzhuang Zhang, Hui-liang Shen

机构 * College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) School of Automotive Studies, Tongji University(同济大学汽车学院) Momoni AI, Gothenburg, Sweden(Momoni AI(瑞典哥德堡)) College of Energy Engineering, Zhejiang University(浙江大学能源工程学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06120 2025-11-11 cs.CV 57%

Bidirectional Image-Event Guided Fusion Framework for Low-Light Image Enhancement

Zhanwen Liu, Huanna Song, Yang Wang, Nan Yang, Weiping Ding, Yisheng An

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19404 2025-11-11 cs.CV 57%

LangBridge: Interpreting Image as a Combination of Language Embeddings

Jiaqi Liao, Yuwei Niu, Fanqing Meng, Hao Li, Changyao Tian, Yinuo Du, Yuwen Xiong, Dianqi Li, Xizhou Zhu, Li Yuan, Jifeng Dai, Yu Cheng

机构 * Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) SenseTime Research(商汤科技研究院) Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) PengCheng Laboratory(鹏城实验室) Chongqing University(重庆大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments The code and weights are open-sourced. Project page: https://curryx-001.github.io/LangBridge.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07081 2025-11-11 cs.RO 50%

HDCNet: A Hybrid Depth Completion Network for Grasping Transparent and Reflective Objects

Guanghu Xie, Mingxu Li, Songwei Wu, Yang Liu, Zongwu Xie, Baoshi Cao, Hong Liu

机构 * State Key Laboratory of Robotics and Systems, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05796 2025-11-11 cs.CR 50%

Securing UAV Communications by Fusing Cross-Layer Fingerprints

Yong Huang, Ruihao Li, Mingyang Chen, Feiyang Zhao, Dalong Zhang, Wanqing Tu

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments To appear in the IEEE Internet of Things Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05726 2025-11-11 cs.LG q-bio.QM 50%

GastroDL-Fusion: A Dual-Modal Deep Learning Framework Integrating Protein-Ligand Complexes and Gene Sequences for Gastrointestinal Disease Drug Discovery

Ziyang Gao, Annie Cheung, Yihao Ou

专题命中 多模态训练与对齐 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏