arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2403.11771 2024-03-19 cs.CV cs.CL 73%

Modality-Agnostic fMRI Decoding of Vision and Language

Mitja Nikolaus, Milad Mozafari, Nicholas Asher, Leila Reddy, Rufin VanRullen

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments To appear at ICLR 2024 workshop on Representational Alignment (Re-Align)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05220 2024-03-11 cs.CV cs.AI cs.LG q-bio.TO 73%

Synthetic Privileged Information Enhances Medical Image Representation Learning

Lucas Farndale, Chris Walsh, Robert Insall, Ke Yuan

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00849 2024-03-11 cs.CL cs.CV 73%

RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback

Tianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun, Tat-Seng Chua

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments Accepted by CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19001 2023-10-31 cs.CV cs.AI cs.LG 73%

Uncovering Prototypical Knowledge for Weakly Open-Vocabulary Semantic Segmentation

Fei Zhang, Tianfei Zhou, Boyang Li, Hao He, Chaofan Ma, Tianjiao Zhang, Jiangchao Yao, Ya Zhang, Yanfeng Wang

专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 14 pages, Accept in NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07704 2023-10-12 cs.CV cs.CL 73%

Ferret: Refer and Ground Anything Anywhere at Any Granularity

Haoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du, Bowen Zhang, Zirui Wang, Liangliang Cao, Shih-Fu Chang, Yinfei Yang

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments 30 pages, 10 figures. Code/Project Website: https://github.com/apple/ml-ferret

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07552 2023-10-12 cs.CV cs.AI 73%

ProtoHPE: Prototype-guided High-frequency Patch Enhancement for Visible-Infrared Person Re-identification

Guiwei Zhang, Yongfei Zhang, Zichang Tan

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00672 2023-10-10 cs.LG cs.CL cs.CV 73%

GeRA: Label-Efficient Geometrically Regularized Alignment

Dustin Klebe, Tal Shnitzer, Mikhail Yurochkin, Leonid Karlinsky, Justin Solomon

专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10729 2023-09-28 cs.CV cs.AI cs.LG 73%

UnICLAM:Contrastive Representation Learning with Adversarial Masking for Unified and Interpretable Medical Vision Question Answering

Chenlu Zhan, Peng Peng, Hongsen Wang, Tao Chen, Hongwei Wang

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10354 2023-08-22 cs.AI cs.CL 73%

Imaginations of WALL-E : Reconstructing Experiences with an Imagination-Inspired Module for Advanced AI Systems

Zeinab Sadat Taghavi, Soroush Gooran, Seyed Arshan Dalili, Hamidreza Amirzadeh, Mohammad Jalal Nematbakhsh, Hossein Sameti

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments 18 pages,

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.09695 2023-08-15 cs.CV cs.GR cs.MM 73%

PersonalTailor: Personalizing 2D Pattern Design from 3D Garment Point Clouds

Sauradip Nag, Anran Qi, Xiatian Zhu, Ariel Shamir

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15765 2023-05-26 cs.CV cs.AI 73%

Language-Guided 3D Object Detection in Point Cloud for Autonomous Driving

Wenhao Cheng, Junbo Yin, Wei Li, Ruigang Yang, Jianbing Shen

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09609 2023-04-20 cs.CV cs.AI 73%

MMDR: A Result Feature Fusion Object Detection Approach for Autonomous System

Wendong Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.02995 2023-03-07 cs.CV cs.CL cs.LG 73%

HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention

Shijie Geng, Jianbo Yuan, Yu Tian, Yuxiao Chen, Yongfeng Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted at ICLR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00902 2023-02-06 cs.LG cs.CL cs.CV 73%

Language Quantized AutoEncoders: Towards Unsupervised Text-Image Alignment

Hao Liu, Wilson Yan, Pieter Abbeel

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments Fixed typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07636 2022-12-06 cs.CV cs.CL cs.LG 73%

EVA: Exploring the Limits of Masked Visual Representation Learning at Scale

Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, Yue Cao

专题命中 多模态训练与对齐 :multi-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments v2: (i) fix / update EVA IN-1K variants results. (ii) add / update EVA-CLIP results. (iii) add Appendix. (iv) release all the code and models at https://github.com/baaivision/EVA

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14777 2022-12-02 cs.CV cs.CL 73%

Alignment-Enriched Tuning for Patch-Level Pre-trained Document Image Models

Lei Wang, Jiabang He, Xing Xu, Ning Liu, Hui Liu

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by AAAI 2023. Code is available at https://github.com/MAEHCM/AET

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.06097 2022-11-14 cs.CV cs.AI 73%

Interactive Context-Aware Network for RGB-T Salient Object Detection

Yuxuan Wang, Feng Dong, Jinchao Zhu

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.13979 2022-08-01 cs.CL cs.AI 73%

Knowing Where and What: Unified Word Block Pretraining for Document Understanding

Song Tao, Zijian Wang, Tiantian Fan, Canjie Luo, Can Huang

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments incomplete experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08387 2022-07-20 cs.CL cs.CV 73%

LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking

Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, Furu Wei

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments ACM Multimedia 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.15119 2022-05-26 cs.CV cs.AI 73%

Aerial Images Meet Crowdsourced Trajectories: A New Approach to Robust Road Extraction

Lingbo Liu, Zewei Yang, Guanbin Li, Kuo Wang, Tianshui Chen, Liang Lin

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments This work has been accepted by IEEE Transactions on Neural Networks and Learning Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.04482 2022-03-31 cs.CV cs.CL 73%

FLAVA: A Foundational Language And Vision Alignment Model

Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, Douwe Kiela

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13285 2022-03-30 cs.SD cs.CV cs.LG eess.AS 73%

Continuous-Time Audiovisual Fusion with Recurrence vs. Attention for In-The-Wild Affect Recognition

Vincent Karas, Mani Kumar Tellamekala, Adria Mallol-Ragolta, Michel Valstar, Björn W. Schuller

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、eess.AS

Comments 10 pages, 1 figures, added references and an overview figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.14973 2021-12-23 cs.CV cs.AI cs.LG cs.RO 73%

MultiPath++: Efficient Information Fusion and Trajectory Aggregation for Behavior Prediction

Balakrishnan Varadarajan, Ahmed Hefny, Avikalp Srivastava, Khaled S. Refaat, Nigamaa Nayakanti, Andre Cornman, Kan Chen, Bertrand Douillard, Chi Pang Lam, Dragomir Anguelov, Benjamin Sapp

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.03331 2021-06-08 cs.CV cs.CL 73%

SelfDoc: Self-Supervised Document Representation Learning

Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, Hongfu Liu

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments To appear in CVPR'2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.11562 2021-01-28 cs.CV cs.CL 73%

Scheduled Sampling in Vision-Language Pretraining with Decoupled Encoder-Decoder Network

Yehao Li, Yingwei Pan, Ting Yao, Jingwen Chen, Tao Mei

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments AAAI 2021; Code is publicly available at: https://github.com/YehLi/TDEN

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.00200 2020-10-01 cs.CV cs.CL cs.LG 73%

HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, Jingjing Liu

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.04464 2020-04-21 cs.CV cs.CL 73%

Relationship-Embedded Representation Learning for Grounding Referring Expressions

Sibei Yang, Guanbin Li, Yizhou Yu

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments This paper is going to appear in TPAMI. Code is available at https://github.com/sibeiyang/sgmn/tree/master/lib/cmrin_models

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05491 2026-06-05 cs.CV cs.RO 72%

Unpaired RGB-Thermal Gaussian-Splatting Using Visual Geometric Transformers

无配对RGB-热成像高斯泼溅使用视觉几何变换器

Jean Cordonnier, Chenghao Xu, Olga Fink, Malcolm Mielle

机构 * Ecole Polytechnique Federale de Lausanne(瑞士联邦理工学院洛桑分校) Schindler EPFL Lab(施耐德EPFL实验室)

专题命中 多模态训练与对齐 :multi-modal(abstract,comments);cross-modal(abstract);分类 cs.CV

AI总结 提出一种无配对RGB-热成像新视角合成框架,利用VGGT估计各模态相机位姿并通过Procrustes对齐,结合多模态3D高斯泼溅实现联合重建,在保持RGB保真度的同时实现热成像视图合成。

Comments Accepted at ICRA 2026's Workshop MM-SpatialAI: Multi-Modal Spatial AI for Robust Navigation and Open-World Understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00086 2026-08-04 cs.CV cs.AI cs.CL cs.LG 版本更新 71%

Hierarchical Pre-Training of Vision Encoders with Large Language Model

基于大语言模型的视觉编码器分层预训练

Eugene Lee, Ting-Yu Chang, Jui-Huang Tsai, Jiajie Diao, Chen-Yi Lee

机构 * University of Cincinnati(辛辛那提大学) National Yang Ming Chiao Tung University(国立阳明交通大学)

专题命中 多模态训练与对齐 :multimodal(abstract,comments);分类 cs.CV、cs.CL、cs.AI;multimodal foundation model(comments)

AI总结 本文提出HIVE框架,通过引入视觉编码器与大语言模型间的分层交叉注意力机制,提升视觉语言对齐,改进特征融合与表征学习,实验表明其在图像分类和多模态任务中表现优异。

Comments 17 pages, 14 figures, accepted to Computer Vision and Pattern Recognition Conference (CVPR) Workshops 2026. 5th MMFM Workshop: What is Next in Multimodal Foundation Models?

Journal ref In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7415-7424) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07407 2026-07-28 cs.LG 版本更新 71%

Emergent Symbolic Structure in Health Foundation Models: Extraction, Alignment, and Cross-Modal Transfer

健康基础模型中的涌现符号结构:提取、对齐与跨模态迁移

Gajendra Katuwal, Advait Koparkar, Salar Abbaspourazad, Anshuman Mishra, Sarvesh Kirthivasan

机构 * Apple(苹果公司)

专题命中 多模态训练与对齐 :cross-modal(title)

AI总结 本文提出一种训练后框架,通过分解冻结嵌入以提取可解释的符号,用于对齐嵌入空间。在PPG和加速度计数据上验证,发现符号能选择性关联健康状况和生理属性,并支持跨模态迁移。

Comments 8 pages, Mechanistic Interpretability Workshop at the 43rd International Conference on Machine Learning, 4 main figures

Journal ref Mechanistic Interpretability Workshop at the 43 rd International Conference on Machine Learning, Seoul, South Korea, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏