arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46294 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4676 篇

2312.11570 2024-03-13 cs.CV 79%

Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model

Shuailei Ma, Chen-Wei Xie, Ying Wei, Siyang Sun, Jiaqi Fan, Xiaoyi Bao, Yuxin Guo, Yun Zheng

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments We find that the statistical information in Figure 2 neglect the statistics for tSOS, so we make corrections. Additionally, we change the statistical samples to those where CLIP misidentify, but prompt tuning identify correctly. At the same time, we also revise some of the descriptions. The changes to the supplementary materials will be updated shortly. arXiv admin note: text overlap with arXiv:2307.06948 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09957 2024-03-13 cs.CV 79%

Self-paced Multi-grained Cross-modal Interaction Modeling for Referring Expression Comprehension

Peihan Miao, Wei Su, Gaoang Wang, Xuewei Li, Xi Li

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by TIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.06295 2024-03-12 cs.CV 79%

A streamlined Approach to Multimodal Few-Shot Class Incremental Learning for Fine-Grained Datasets

Thang Doan, Sima Behpour, Xin Li, Wenbin He, Liang Gou, Liu Ren

专题命中 图文多模态 :multimodal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05141 2024-03-11 cs.CV 79%

Med3DInsight: Enhancing 3D Medical Image Understanding with 2D Multi-Modal Large Language Models

Qiuhui Chen, Huping Ye, Yi Hong

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08825 2024-03-11 cs.CV 79%

From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Dongsheng Jiang, Yuchen Liu, Songlin Liu, Jin'e Zhao, Hao Zhang, Zhen Gao, Xiaopeng Zhang, Jin Li, Hongkai Xiong

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.12105 2024-03-06 cs.LG cs.CL cs.SE 79%

Mass-Producing Failures of Multimodal Systems with Language Models

Shengbang Tong, Erik Jones, Jacob Steinhardt

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.00249 2024-03-04 cs.CV 79%

Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training

Haowei Liu, Yaya Shi, Haiyang Xu, Chunfeng Yuan, Qinghao Ye, Chenliang Li, Ming Yan, Ji Zhang, Fei Huang, Bing Li, Weiming Hu

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17680 2024-02-28 cs.CV 79%

MCF-VC: Mitigate Catastrophic Forgetting in Class-Incremental Learning for Multimodal Video Captioning

Huiyu Xiong, Lanxiao Wang, Heqian Qiu, Taijin Zhao, Benliu Qiu, Hongliang Li

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11325 2024-02-28 cs.CV 79%

ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models

Zhenghang Yuan, Zhitong Xiong, Lichao Mou, Xiao Xiang Zhu

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05496 2024-02-27 cs.CL 79%

Exploiting Pseudo Image Captions for Multimodal Summarization

Chaoya Jiang, Rui Xie, Wei Ye, Jinan Sun, Shikun Zhang

专题命中 图文多模态 :multimodal(title);cross-modal(abstract);分类 cs.CL

Comments Accepted at ACL2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14767 2024-02-23 cs.CV 79%

DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models

Yuhang Cao, Pan Zhang, Xiaoyi Dong, Dahua Lin, Jiaqi Wang

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06659 2024-02-21 cs.CL 79%

WisdoM: Improving Multimodal Sentiment Analysis by Fusing Contextual World Knowledge

Wenbin Wang, Liang Ding, Li Shen, Yong Luo, Han Hu, Dacheng Tao

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12048 2024-02-20 cs.CL 79%

Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models

Didi Zhu, Zhongyi Sun, Zexi Li, Tao Shen, Ke Yan, Shouhong Ding, Kun Kuang, Chao Wu

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08670 2024-02-14 cs.AI 79%

Rec-GPT4V: Multimodal Recommendation with Large Vision-Language Models

Yuqing Liu, Yu Wang, Lichao Sun, Philip S. Yu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06092 2024-02-12 cs.CV cs.RO 79%

CLIP-Loc: Multi-modal Landmark Association for Global Localization in Object-based Maps

Shigemichi Matsuzaki, Takuma Sugino, Kazuhito Tanaka, Zijun Sha, Shintaro Nakaoka, Shintaro Yoshizawa, Kazuhiro Shintani

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 7 pages, 7 figures. Accepted to IEEE International Conference on Robotics and Automation (ICRA) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.05472 2024-02-09 cs.CV 79%

Question Aware Vision Transformer for Multimodal Reasoning

Roy Ganz, Yair Kittenplon, Aviad Aberdam, Elad Ben Avraham, Oren Nuriel, Shai Mazor, Ron Litman

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17862 2024-02-01 cs.CV 79%

Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis

Jianing Li, Xi Nan, Ming Lu, Li Du, Shanghang Zhang

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 15 pages,version 1

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.13307 2024-01-25 cs.CV 79%

ChatterBox: Multi-round Multimodal Referring and Grounding

Yunjie Tian, Tianren Ma, Lingxi Xie, Jihao Qiu, Xi Tang, Yuan Zhang, Jianbin Jiao, Qi Tian, Qixiang Ye

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 17 pages, 6 tables, 9 figurs. Code, data, and model are available at: https://github.com/sunsmarterjie/ChatterBox

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03851 2024-01-09 cs.CV q-bio.NC 79%

Aligned with LLM: a new multi-modal training paradigm for encoding fMRI activity in visual cortex

Shuxiao Ma, Linyuan Wang, Senbao Hou, Bin Yan

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01911 2024-01-05 cs.CV cs.LG 79%

Backdoor Attack on Unpaired Medical Image-Text Foundation Models: A Pilot Study on MedCLIP

Ruinan Jin, Chun-Yin Huang, Chenyu You, Xiaoxiao Li

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Paper Accepted at the 2nd IEEE Conference on Secure and Trustworthy Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01076 2024-01-04 cs.CL 79%

DialCLIP: Empowering CLIP as Multi-Modal Dialog Retriever

Zhichao Yin, Binyuan Hui, Min Yang, Fei Huang, Yongbin Li

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CL

Comments ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.16602 2023-12-29 cs.CV 79%

Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Jiaxing Huang, Jingyi Zhang, Kai Jiang, Han Qiu, Shijian Lu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14465 2023-12-25 cs.CV 79%

FM-OV3D: Foundation Model-based Cross-modal Knowledge Blending for Open-Vocabulary 3D Detection

Dongmei Zhang, Chang Li, Ray Zhang, Shenghao Xie, Wei Xue, Xiaodong Xie, Shanghang Zhang

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by AAAI 2024. Code will be released at https://github.com/dmzhang0425/FM-OV3D.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08548 2023-12-15 cs.CV 79%

EVP: Enhanced Visual Perception using Inverse Multi-Attentive Feature Refinement and Regularized Image-Text Alignment

Mykola Lavreniuk, Shariq Farooq Bhat, Matthias Müller, Peter Wonka

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03777 2023-12-11 cs.CV 79%

On the Robustness of Large Multimodal Models Against Image Adversarial Attacks

Xuanming Cui, Alejandro Aparcedo, Young Kyun Jang, Ser-Nam Lim

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01191 2023-12-05 cs.CV 79%

Bootstrapping Interactive Image-Text Alignment for Remote Sensing Image Captioning

Cong Yang, Zuchao Li, Lefei Zhang

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00438 2023-12-04 cs.CV 79%

Dolphins: Multimodal Language Model for Driving

Yingzi Ma, Yulong Cao, Jiachen Sun, Marco Pavone, Chaowei Xiao

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments The project page is available at https://vlm-driver.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17583 2023-11-30 cs.CV 79%

CLIPC8: Face liveness detection algorithm based on image-text pairs and contrastive learning

Xu Liu, Shu Zhou, Yurong Song, Wenzhe Luo, Xin Zhang

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13847 2023-11-29 cs.CV cs.IT eess.IV math.IT 79%

Perceptual Image Compression with Cooperative Cross-Modal Side Information

Shiyu Qin, Bin Chen, Yujun Huang, Baoyi An, Tao Dai, Shu-Tao Xia

专题命中 图文多模态 :cross-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05425 2023-11-10 cs.CV 79%

Active Mining Sample Pair Semantics for Image-text Matching

Yongfeng Chena, Jin Liua, Zhijing Yang, Ruihan Chena, Junpeng Tan

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏