arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6918 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6918 篇

2503.05933 2025-11-18 eess.IV cs.CV 57%

Beyond H&E: Unlocking Pathological Insights with Polarization Imaging

Yao Du, Jiaxin Zhuang, Xiaoyu Zheng, Jing Cong, Limei Guo, Chao He, Lin Luo, Xiaomeng Li

机构 * The Hong Kong University of Science and Technology(香港科技大学) Beijing Institute of Collaborative Innovation(北京协同创新研究院) Peking University Health Science Center, Peking University Third Hospital(北京大学人民医院) University of Oxford(牛津大学) Peking University(北京大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Accepted as a regular paper at IEEE BIBM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11422 2025-11-17 cs.CV 57%

Shrinking the Teacher: An Adaptive Teaching Paradigm for Asymmetric EEG-Vision Alignment

Lukun Wu, Jie Li, Ziqi Ren, Kaifan Zhang, Xinbo Gao

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 21pages,12 figures,published to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10560 2025-11-17 cs.CV 57%

OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer

Haosong Peng, Hao Li, Yalun Dai, Yushi Lan, Yihang Luo, Tianyu Qi, Zhengshen Zhang, Yufeng Zhan, Junfei Zhang, Wenchao Xu, Ziwei Liu

机构 * HKUST(香港科技大学) NTU(国立台湾大学) SYSU(南方科技大学) NUS(国立新加坡大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments Project Page: https://livioni.github.io/OmniVGGT-official/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15888 2025-11-17 cs.CV 57%

MS-Occ: Multi-Stage LiDAR-Camera Fusion for 3D Semantic Occupancy Prediction

Zhiqiang Wei, Lianqing Zheng, Jianan Liu, Tao Huang, Qing-Long Han, Wenwen Zhang, Fengdeng Zhang

机构 * School of Optical-Electrical and Computer Engineering, University of Shanghai for Science and Technology(光学电子与计算机工程学院,上海科学技术大学) School of Automotive Studies, Tongji University(汽车学院,同济大学) Momoni AI College of Science and Engineering, James Cook University(科学与工程学院,詹姆斯库克大学) School of Engineering, Swinburne University of Technology(工程学院,斯威本技术大学) School of Electrical and Electronic Engineering, Nanyang Technological University(电气与电子工程学院,南洋理工大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10211 2025-11-14 cs.CV 57%

HeatV2X: Scalable Heterogeneous Collaborative Perception via Efficient Alignment and Interaction

Yueran Zhao, Zhang Zhang, Chao Sun, Tianze Wang, Chao Yue, Nuoran Li

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10087 2025-11-14 cs.RO cs.AI cs.LG 57%

Opinion: Towards Unified Expressive Policy Optimization for Robust Robot Learning

Haidong Huang, Haiyue Zhu. Jiayu Song, Xixin Zhao, Yaohua Zhou, Jiayi Zhang, Yuze Zhai, Xiaocong Li

机构 * Eastern Institute of Technology(东部技术研究所) University of Nottingham(诺丁汉大学) SIMTech, Agency for Science, Technology and Research (A*STAR)(SIMTech,科技研究局(A*STAR)) Southern University of Science and Technology(南方科技大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025 Workshop on Embodied World Models for Decision Making

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08903 2025-11-14 cs.CV 57%

LLM-Guided Probabilistic Fusion for Label-Efficient Document Layout Analysis

Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

机构 * Department of Computer Science, Iowa State University(计算机科学系,爱荷华州立大学) Department of Civil, Construction & Environmental Engineering, Iowa State University(土木、建设与环境工程系,爱荷华州立大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00598 2025-11-12 cs.CV 57%

DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation

Boyi Li, Ce Zhang, Richard M. Timmerman, Wenxuan Bao

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03565 2025-11-12 cs.CV 57%

Bridged Semantic Alignment for Zero-shot 3D Medical Image Diagnosis

Haoran Lai, Zihang Jiang, Qingsong Yao, Rongsheng Wang, Zhiyang He, Xiaodong Tao, Weifu Lv, Wei Wei, S. Kevin Zhou

机构 * University of Science and Technology of China(中国科学技术大学) Suzhou Institute for Advanced Research(苏州先进研究所) Stanford University(斯坦福大学) iFlytek Co. Ltd.(iFlytek公司) The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine, USTC(中国科学技术大学第一附属医院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06120 2025-11-11 cs.CV 57%

Bidirectional Image-Event Guided Fusion Framework for Low-Light Image Enhancement

Zhanwen Liu, Huanna Song, Yang Wang, Nan Yang, Weiping Ding, Yisheng An

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19404 2025-11-11 cs.CV 57%

LangBridge: Interpreting Image as a Combination of Language Embeddings

Jiaqi Liao, Yuwei Niu, Fanqing Meng, Hao Li, Changyao Tian, Yinuo Du, Yuwen Xiong, Dianqi Li, Xizhou Zhu, Li Yuan, Jifeng Dai, Yu Cheng

机构 * Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) SenseTime Research(商汤科技研究院) Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) PengCheng Laboratory(鹏城实验室) Chongqing University(重庆大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments The code and weights are open-sourced. Project page: https://curryx-001.github.io/LangBridge.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05044 2025-11-10 cs.CV 57%

Medical Referring Image Segmentation via Next-Token Mask Prediction

Xinyu Chen, Yiran Wang, Gaoyang Pang, Jiafu Hao, Chentao Yue, Luping Zhou, Yonghui Li

机构 * School of Electrical and Computer Engineering, University of Sydney(悉尼大学电气与计算机工程学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments This work has been submitted to the IEEE Transactions on Medical Imaging for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04790 2025-11-10 cs.LG cs.AI stat.ML 57%

Causal Structure and Representation Learning with Biomedical Applications

Caroline Uhler, Jiaqi Zhang

机构 * Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology(电气工程与计算机科学系,麻省理工学院)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.AI

Comments This article has successfully completed peer review and will appear in the Proceedings of the International Congress of Mathematicians 2026. Both authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15132 2025-11-06 cs.CV eess.IV 57%

FusionRF: High-Fidelity Satellite Neural Radiance Fields from Multispectral and Panchromatic Acquisitions

Michael Sprintson, Rama Chellappa, Cheng Peng

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02685 2025-11-05 cs.CV 57%

Modality-Transition Representation Learning for Visible-Infrared Person Re-Identification

Chao Yuan, Zanwu Liu, Guiwei Zhang, Haoxuan Xu, Yujian Zhao, Guanglin Niu, Bo Li

机构 * Beihang University(北航大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26466 2025-11-04 cs.CV cs.LG 57%

Representation-Level Counterfactual Calibration for Debiased Zero-Shot Recognition

Pei Peng, MingKun Xie, Hang Hao, Tong Jin, ShengJun Huang

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19739 2025-11-04 cs.CV 57%

FUSE: Label-Free Image-Event Joint Monocular Depth Estimation via Frequency-Decoupled Alignment and Degradation-Robust Fusion

Pihai Sun, Junjun Jiang, Yuanqi Yao, Youyu Chen, Wenbo Zhao, Kui Jiang, Xianming Liu

机构 * Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Zhengzhou Research Institute, Harbin Institute of Technology(郑州研究院,哈尔滨工业大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments [IROS 2025, camera ready version]: 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24385 2025-10-30 cs.CV 57%

When are radiology reports useful for training medical image classifiers?

Herman Bergström, Zhongqi Yue, Fredrik D. Johansson

机构 * Department of Computer Science & Engineering, Chalmers University of Technology and University of Gothenburg(计算机科学与工程系,楚姆勒斯技术大学和哥德堡大学)

专题命中 多模态训练与对齐 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19638 2025-10-30 cs.CV 57%

HF-VTON: High-Fidelity Virtual Try-On via Consistent Geometric and Semantic Alignment

Ming Meng, Qi Dong, Jiajie Li, Zhe Zhu, Xingyu Wang, Zhaoxin Fan, Wei Zhao, Wenjun Wu

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments After the publication of the paper, we discovered some significant errors/omissions that need to be corrected and improved

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11245 2025-10-30 cs.CV 57%

L2RSI: Cross-view LiDAR-based Place Recognition for Large-scale Urban Scenes via Remote Sensing Imagery

Ziwei Shi, Xiaoran Zhang, Wenjing Xu, Yan Xia, Yu Zang, Siqi Shen, Cheng Wang

机构 * Fujian Key Laboratory of Sensing and Computing for Smart Cities, Xiamen University, China(福建智能城市感知与计算重点实验室,厦门大学,中国) Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, China(多媒体可信感知与高效计算重点实验室,中华人民共和国教育部,厦门大学,中国) University of Science and Technology of China, China(中国科学技术大学,中国)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments 17 pages, 7 figures, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24551 2025-10-29 cs.AI 57%

Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives

Gang Chen, Changshuo Liu, Gene Anne Ooi, Marcus Tan, Zhongle Xie, Jianwei Yin, James Wei Luen Yip, Wenqiao Zhang, Jiaqi Zhu, Beng Chin Ooi

机构 * College of Computer Science and Technology, Zhejiang University, Hangzhou 310027, China(浙江大学计算机科学与技术学院) College of Software Technology, Zhejiang University, Ningbo 315100, China(浙江大学软件技术学院) School of Computing, National University of Singapore, Singapore 117417(新加坡国立大学计算机学院) Singapore General Hospital, Singapore 169608(新加坡中央医院) National University Hospital, Singapore 119074(新加坡国立医院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14643 2025-10-29 cs.CV 57%

Multispectral State-Space Feature Fusion: Bridging Shared and Cross-Parametric Interactions for Object Detection

Jifeng Shen, Haibo Zhan, Shaohua Dong, Xin Zuo, Wankou Yang, Haibin Ling

机构 * School of Electrical and Information Engineering, Jiangsu University, Zhenjiang, 212013, China(江苏大学电气与信息工程学院) Department of Computer Science and Engineering, University of North Texas, Denton, TX 76207, USA(德克萨斯大学北卡罗来纳分校计算机科学与工程系) School of Computer Science and Engineering, Jiangsu University of Science and Technology, Zhenjiang, 212003, China(江苏科技大学计算机科学与工程学院) School of Automation, Southeast University, Nanjing, 210096, China(东南大学自动化学院) Bodhi Intelligence Lab, Department of Artificial Intelligence, Westlake University, Hangzhou, Zhejiang 310030, China(西湖大学人工智能研究院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments submitted on 30/4/2025, Accepted by Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24214 2025-10-29 cs.CV 57%

SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs

Jinhong Deng, Wen Li, Joey Tianyi Zhou, Yang He

机构 * University of Electronic Science and Technology of China(电子科技大学) Shenzhen Institute for Advanced Study(深圳先进研究 institute) CFAR, Agency for Science, Technology and Research (A*STAR)(科技研究局(A*STAR)认知与人工智能研究中心) IHPC, Agency for Science, Technology and Research (A*STAR)(科技研究局(A*STAR)人工智能中心)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00034 2025-10-29 cs.RO cs.CV 57%

GaussianFusion: Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous Driving

Shuai Liu, Quanmin Liang, Zefeng Li, Boyang Li, Kai Huang

机构 * School of Computer Science and Engineering, Sun Yat-sen University(计算机科学与工程学院,中山大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted at NeurIPS2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03201 2025-10-28 cs.CV 57%

AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding

Yidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao, Nicu Sebe

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Trento(特伦托大学)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22301 2025-10-28 cs.LG cs.AI 57%

AnyECG-Lab: An Exploration Study of Fine-tuning an ECG Foundation Model to Estimate Laboratory Values from Single-Lead ECG Signals

Yujie Xiao, Gongzhen Tang, Wenhui Liu, Jun Li, Guangkun Nie, Zhuoran Kan, Deyun Zhang, Qinghao Zhao, Shenda Hong

机构 * Institute of Medical Technology, Peking University Health Science Center(北京大学人民医院医学技术研究所) National Institute of Health Data Science, Peking University(北京大学国家健康数据科学研究院) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) HeartVoice Medical Technology(心声医疗技术) Department of Cardiology, Peking University People’s Hospital(北京大学人民医院心内科) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) State Key Laboratory of Vascular Homeostasis and Remodeling, NHC Key Laboratory of Cardiovascular Molecular Biology and Regulatory Peptides, Peking University(国家心血管病分子生物学与调节肽重点实验室,北京大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16188 2025-10-28 cs.CV 57%

Think or Not Think: A Study of Explicit Thinking in Rule-Based Visual Reinforcement Fine-Tuning

Ming Li, Jike Zhong, Shitian Zhao, Yuxiang Lai, Haoquan Zhang, Wang Bill Zhu, Kaipeng Zhang

机构 * Shanghai AI Laboratory(上海人工智能实验室) University of Southern California(南加州大学) Emory University(埃默里大学) Chinese University of Hong Kong(香港中文大学)

专题命中 多模态训练与对齐 :MLLM(abstract);分类 cs.CV

Comments Neurips 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21553 2025-10-27 cs.CL cs.LG 57%

Document Understanding, Measurement, and Manipulation Using Category Theory

Jared Claypoole, Yunye Gong, Noson S. Yanofsky, Ajay Divakaran

机构 * SRI International(SRI国际研究院) Brooklyn College(布鲁克林学院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20345 2025-10-24 cs.AI 57%

LLM-empowered knowledge graph construction: A survey

Haonan Bian

机构 * Xidian University(西安电子科技大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19520 2025-10-23 cs.MM 57%

CDI-DTI: A Strong Cross-domain Interpretable Drug-Target Interaction Prediction Framework Based on Multi-Strategy Fusion

Xiangyu Li, Haojie Yang, Kaimiao Hu, Runzhi Wu, Liangliang Liu, Ran Su

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏