arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-11 至 2025-11-11 共收录 105 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 16 篇

2502.00801 2025-11-11 cs.CV cs.AI cs.RO 73%

Environment-Driven Online LiDAR-Camera Extrinsic Calibration

Zhiwei Huang, Jiaqi Li, Hongbo Zhao, Xiao Ma, Ping Zhong, Xiaohu Zhou, Wei Ye, Rui Fan

机构 * Department of Control Science & Engineering, the College of Electronic & Information Engineering, Tongji University(控制科学与工程系,电子与信息工程学院,同济大学) School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学) Beijing Institute of Aerospace Control Devices(北京航天控制器件研究所) Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06225 2025-11-11 cs.CV 70%

MoRA: Missing Modality Low-Rank Adaptation for Visual Recognition

Shu Zhao, Nilesh Ahuja, Tan Yu, Tianyi Shen, Vijaykrishnan Narayanan

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Intel(英特尔) NVIDIA(英伟达)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06836 2025-11-11 cs.CV cs.AI 62%

NeuroBridge: Bio-Inspired Self-Supervised EEG-to-Image Decoding via Cognitive Priors and Bidirectional Semantic Alignment

Wenjiang Zhang, Sifeng Wang, Yuwei Su, Xinyu Li, Chen Zhang, Suyu Zhong

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19856 2025-11-11 cs.CV cs.AI 62%

RaGS: Unleashing 3D Gaussian Splatting from 4D Radar and Monocular Cues for 3D Object Detection

Xiaokai Bai, Chenxu Zhou, Lianqing Zheng, Si-Yuan Cao, Jianan Liu, Xiaohan Zhang, Yiming Li, Zhengzhuang Zhang, Hui-liang Shen

机构 * College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) School of Automotive Studies, Tongji University(同济大学汽车学院) Momoni AI, Gothenburg, Sweden(Momoni AI(瑞典哥德堡)) College of Energy Engineering, Zhejiang University(浙江大学能源工程学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06120 2025-11-11 cs.CV 57%

Bidirectional Image-Event Guided Fusion Framework for Low-Light Image Enhancement

Zhanwen Liu, Huanna Song, Yang Wang, Nan Yang, Weiping Ding, Yisheng An

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19404 2025-11-11 cs.CV 57%

LangBridge: Interpreting Image as a Combination of Language Embeddings

Jiaqi Liao, Yuwei Niu, Fanqing Meng, Hao Li, Changyao Tian, Yinuo Du, Yuwen Xiong, Dianqi Li, Xizhou Zhu, Li Yuan, Jifeng Dai, Yu Cheng

机构 * Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) SenseTime Research(商汤科技研究院) Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) PengCheng Laboratory(鹏城实验室) Chongqing University(重庆大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments The code and weights are open-sourced. Project page: https://curryx-001.github.io/LangBridge.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07081 2025-11-11 cs.RO 50%

HDCNet: A Hybrid Depth Completion Network for Grasping Transparent and Reflective Objects

Guanghu Xie, Mingxu Li, Songwei Wu, Yang Liu, Zongwu Xie, Baoshi Cao, Hong Liu

机构 * State Key Laboratory of Robotics and Systems, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学)

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05796 2025-11-11 cs.CR 50%

Securing UAV Communications by Fusing Cross-Layer Fingerprints

Yong Huang, Ruihao Li, Mingyang Chen, Feiyang Zhao, Dalong Zhang, Wanqing Tu

专题命中 多模态训练与对齐 :multimodal(abstract)

Comments To appear in the IEEE Internet of Things Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05726 2025-11-11 cs.LG q-bio.QM 50%

GastroDL-Fusion: A Dual-Modal Deep Learning Framework Integrating Protein-Ligand Complexes and Gene Sequences for Gastrointestinal Disease Drug Discovery

Ziyang Gao, Annie Cheung, Yihao Ou

专题命中 多模态训练与对齐 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 其他多模态 6 篇

2511.05952 2025-11-11 cs.HC cs.CV cs.MM 84%

Pinching Visuo-haptic Display: Investigating Cross-Modal Effects of Visual Textures on Electrostatic Cloth Tactile Sensations

Takekazu Kitagishi, Chun-Wei Ooi, Yuichi Hiroi, Jun Rekimoto

机构 * The University of Tokyo(东京大学) ZOZO Research(ZOZO研究) Cluster Metaverse Lab(集群元宇宙实验室) Sony CSL Kyoto(索尼 CSL京都)

专题命中 其他多模态 :cross-modal(title,abstract);multimodal(abstract,comments);分类 cs.CV、cs.MM

Comments 10 pages, 8 figures, 3 tables. Presented at ACM International Conference on Multimodal Interaction (ICMI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06749 2025-11-11 cs.RO cs.CV 79%

Semi-distributed Cross-modal Air-Ground Relative Localization

Weining Lu, Deer Bin, Lian Ma, Ming Ma, Zhihao Ma, Xiangyang Chen, Longfei Wang, Yixiao Feng, Zhouxian Jiang, Yongliang Shi, Bin Liang

机构 * Beijng National Research Center for Information Science and Technology(北京国家信息科学与技术研究中心) Qiyuan Lab(启元实验室) JiangHuai Advanced Technology Center(江淮先进技术中心)

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CV

Comments 7 pages, 3 figures. Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07004 2025-11-11 cs.CV cs.HC 57%

Exploring the "Great Unseen" in Medieval Manuscripts: Instance-Level Labeling of Legacy Image Collections with Zero-Shot Models

Christofer Meinecke, Estelle Guéville, David Joseph Wrisley

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06744 2025-11-11 cs.CV 57%

PointCubeNet: 3D Part-level Reasoning with 3x3x3 Point Cloud Blocks

Da-Yeong Kim, Yeong-Jun Cho

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06297 2025-11-11 cs.HC cs.AI 57%

Decomate: Leveraging Generative Models for Co-Creative SVG Animation

Jihyeon Park, Jiyoon Myung, Seone Shin, Jungki Son, Joohyung Han

机构 * MODULABS

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted at the 1st Workshop on Generative and Protective AI for Content Creation (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06256 2025-11-11 cs.CV 57%

VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving

Ruifei Zhang, Wei Zhang, Xiao Tan, Sibei Yang, Xiang Wan, Xiaonan Luo, Guanbin Li

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院) Sun Yat-sen University(中山大学) Baidu Inc.(百度公司) Guilin University of Electronic Technology(桂林电子科技大学) Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)

专题命中 其他多模态 :MLLM(abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏