arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6903 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6903 篇

2304.14204 2023-04-28 cs.AI cs.CV 84%

Towards Medical Artificial General Intelligence via Knowledge-Enhanced Multimodal Pretraining

Bingqian Lin, Zicong Chen, Mingjie Li, Haokun Lin, Hang Xu, Yi Zhu, Jianzhuang Liu, Wenjia Cai, Lei Yang, Shen Zhao, Chenfei Wu, Ling Chen, Xiaojun Chang, Yi Yang, Lei Xing, Xiaodan Liang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Project page: https://github.com/chenzcv7/MOTOR

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05554 2023-04-13 cs.CV cs.AI 84%

Learning Transferable Pedestrian Representation from Multimodal Information Supervision

Liping Bao, Longhui Wei, Xiaoyu Qiu, Wengang Zhou, Houqiang Li, Qi Tian

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02407 2023-04-06 cs.CV cs.AI cs.LG 84%

Explaining Multimodal Data Fusion: Occlusion Analysis for Wilderness Mapping

Burak Ekim, Michael Schmitt

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02131 2023-03-16 cs.CV cs.CL cs.LG 84%

Masked Vision and Language Modeling for Multi-modal Representation Learning

Gukyeong Kwon, Zhaowei Cai, Avinash Ravichandran, Erhan Bas, Rahul Bhotika, Stefano Soatto

专题命中 多模态训练与对齐 :multi-modal(title);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments International Conference on Learning Representations (ICLR) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.05952 2023-03-13 cs.LG cs.AI cs.CV 84%

Understanding and Constructing Latent Modality Structures in Multi-modal Representation Learning

Qian Jiang, Changyou Chen, Han Zhao, Liqun Chen, Qing Ping, Son Dinh Tran, Yi Xu, Belinda Zeng, Trishul Chilimbi

专题命中 多模态训练与对齐 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 14 pages, 8 figure, CVPR 2023 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05543 2023-02-17 cs.CV cs.MM 84%

Towards Reliable Image Outpainting: Learning Structure-Aware Multimodal Fusion with Depth Guidance

Lei Zhang, Kang Liao, Chunyu Lin, Yao Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.15824 2023-01-31 cs.MM cs.AI 84%

Improving the Modality Representation with Multi-View Contrastive Learning for Multimodal Sentiment Analysis

Peipei Liu, Xin Zheng, Hong Li, Jie Liu, Yimo Ren, Hongsong Zhu, Limin Sun

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11362 2023-01-30 cs.CV cs.CL 84%

Improving Cross-modal Alignment for Text-Guided Image Inpainting

Yucheng Zhou, Guodong Long

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments EACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00678 2022-12-02 cs.CL cs.CV cs.LG 84%

Adapted Multimodal BERT with Layer-wise Fusion for Sentiment Analysis

Odysseas S. Chlapanis, Georgios Paraskevopoulos, Alexandros Potamianos

专题命中 多模态训练与对齐 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.10916 2022-11-29 cs.CV cs.CL cs.LG 84%

Hierachical Delta-Attention Method for Multimodal Fusion

Kunjal Panchal

专题命中 多模态训练与对齐 :multimodal(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Comments Need to update the results

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00526 2022-11-02 cs.CL cs.AI 84%

Leveraging Graph-based Cross-modal Information Fusion for Neural Sign Language Translation

Jiangbin Zheng, Siyuan Li, Cheng Tan, Chong Wu, Yidong Chen, Stan Z. Li

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09227 2022-04-21 cs.CL cs.SD eess.AS 84%

Cross-stitched Multi-modal Encoders

Karan Singla, Daniel Pressel, Ryan Price, Bhargav Srinivas Chinnari, Yeon-Jun Kim, Srinivas Bangalore

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.05784 2022-03-14 eess.IV cs.AI cs.CV 84%

AI-enabled Automatic Multimodal Fusion of Cone-Beam CT and Intraoral Scans for Intelligent 3D Tooth-Bone Reconstruction and Clinical Applications

Jin Hao, Jiaxiang Liu, Jin Li, Wei Pan, Ruizhe Chen, Huimin Xiong, Kaiwei Sun, Hangzheng Lin, Wanlu Liu, Wanghui Ding, Jianfei Yang, Haoji Hu, Yueling Zhang, Yang Feng, Zeyu Zhao, Huikai Wu, Youyi Zheng, Bing Fang, Zuozhu Liu, Zhihe Zhao

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 30 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.14457 2022-01-06 cs.CL cs.AI cs.LG 84%

Towards a Multi-modal, Multi-task Learning based Pre-training Framework for Document Representation Learning

Subhojeet Pramanik, Shashank Mujumdar, Hima Patel

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.13731 2021-08-11 cs.CV cs.AI 84%

UIBert: Learning Generic Multimodal Representations for UI Understanding

Chongyang Bai, Xiaoxue Zang, Ying Xu, Srinivas Sunkara, Abhinav Rastogi, Jindong Chen, Blaise Aguera y Arcas

专题命中 多模态训练与对齐 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI

Comments 8 pages, IJCAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.11672 2021-05-26 cs.CL cs.CV cs.LG 84%

ViBERTgrid: A Jointly Trained Multi-Modal 2D Document Representation for Key Information Extraction from Documents

Weihong Lin, Qifang Gao, Lei Sun, Zhuoyao Zhong, Kai Hu, Qin Ren, Qiang Huo

专题命中 多模态训练与对齐 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments To be published at ICDAR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03435 2021-04-09 cs.CV cs.AI 84%

Multimodal Fusion Refiner Networks

Sethuraman Sankaran, David Yang, Ser-Nam Lim

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.04727 2021-01-14 cs.CL cs.AI 84%

Latent Alignment of Procedural Concepts in Multimodal Recipes

Hossein Rajaby Faghihi, Roshanak Mirzaee, Sudarshan Paliwal, Parisa Kordjamshidi

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments Published in ALVR 2020, a workshop in ACL 2020

Journal ref Proceedings of the First Workshop on Advances in Language and Vision Research 2020 (26-31)

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.00514 2020-10-02 cs.CV cs.CL 84%

Referring Image Segmentation via Cross-Modal Progressive Comprehension

Shaofei Huang, Tianrui Hui, Si Liu, Guanbin Li, Yunchao Wei, Jizhong Han, Luoqi Liu, Bo Li

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by CVPR 2020. Code is available at https://github.com/spyflying/CMPC-Refseg

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.11740 2020-07-21 cs.CV cs.CL cs.LG 84%

UNITER: UNiversal Image-TExt Representation Learning

Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, Jingjing Liu

专题命中 多模态训练与对齐 :image-text(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL

Comments ECCV 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.07407 2018-11-20 cs.CV cs.AI cs.LG 84%

Multimodal Densenet

Faisal Mahmood, Ziyun Yang, Thomas Ashley, Nicholas J. Durr

专题命中 多模态训练与对齐 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.03414 2018-10-09 cs.CV cs.MM 84%

Dense Multimodal Fusion for Hierarchically Joint Representation

Di Hu, Feiping Nie, Xuelong Li

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.03920 2018-08-14 cs.LG cs.AI cs.CL cs.NE stat.ML 84%

Multimodal Language Analysis with Recurrent Multistage Fusion

Paul Pu Liang, Ziyin Liu, Amir Zadeh, Louis-Philippe Morency

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.09534 2016-11-30 cs.CV cs.CL 84%

Is a picture worth a thousand words? A Deep Multi-Modal Fusion Architecture for Product Classification in e-commerce

Tom Zahavy, Alessandro Magnani, Abhinandan Krishnan, Shie Mannor

专题命中 多模态训练与对齐 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20198 2026-02-03 cs.CV 84%

A Survey of Token Compression for Efficient Multimodal Large Language Models

多模态大语言模型高效性中的标记压缩综述

Kele Shao, Keda Tao, Kejia Zhang, Sicheng Feng, Mu Cai, Yuzhang Shang, Haoxuan You, Can Qin, Yang Sui, Huan Wang

机构 * Zhejiang University(浙江大学) Westlake University(西湖大学) Xiamen University(厦门大学) National University of Singapore(新加坡国立大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) University of Central Florida(佛罗里达大学) Salesforce AI Research(Salesforce AI研究) Rice University(德克萨斯大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文综述了多模态大语言模型中标记压缩技术,分类讨论了图像、视频和音频三种模态的压缩方法及其机制,旨在推动该领域的发展。

Comments For ongoing updates and to track the latest advances in this promising area, we maintain a public repository: https://github.com/cokeshao/Awesome-Multimodal-Token-Compression

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19491 2024-07-30 cs.CV 84%

Multi-modal Crowd Counting via Modal Emulation

Chenhao Wang, Xiaopeng Hong, Zhiheng Ma, Yupeng Wei, Yabin Wang, Xiaopeng Fan

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments This is the preprint version of the paper to appear in BMVC 2024. Please cite the final published version. Code is available at https://github.com/Mr-Monday/Multi-modal-Crowd-Counting-via-Modal-Emulation

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09513 2024-03-15 cs.CR cs.AI 84%

AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting

Yu Wang, Xiaogeng Liu, Yu Li, Muhao Chen, Chaowei Xiao

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Multimodal Large Language Models Defense, 25 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12898 2023-08-28 cs.MM cs.AI cs.CL cs.CV 84%

Can Linguistic Knowledge Improve Multimodal Alignment in Vision-Language Pretraining?

Fei Wang, Liang Ding, Jun Rao, Ye Liu, Li Shen, Changxing Ding

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments [TL;DR] we design and release the SNARE, the first large-scale multimodal alignment probing benchmark for current vision-language pretrained models

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11991 2026-05-04 cs.CV cs.AI cs.CL 83%

VGR: Visual Grounded Reasoning

VGR:视觉基础推理

Jiacong Wang, Zijian Kang, Haochen Wang, Haiyong Jiang, Jiawen Li, Bohong Wu, Ya Wang, Jiao Ran, Xiao Liang, Chao Feng, Jun Xiao

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ByteDance Inc.(字节跳动公司)

专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出VGR,一种增强视觉感知的多模态大语言模型,通过图像区域检测与回放提升多模态推理能力,在多个基准测试中表现优异。

Comments 9 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19971 2026-08-21 cs.CL 新提交 83%

Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction

基于迭代代理修正的鲁棒性不完整多模态情感分析

Zhifa Geng, Subin Huang, Hao Guo, Junjie Chen, Sanmin Liu, Chao Kong

机构 * Renmin University of China(中国人民大学) Anhui Polytechnic University(安徽工程大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 针对不完整多模态情感分析中一次性代理初始化粗糙的问题,提出迭代代理修正框架,在MOSI等数据集上实现了优于基线的鲁棒情感预测。

Comments Accepted to SEKE 2026. 6 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏