arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46294 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4676 篇

2406.17047 2024-06-26 cs.CV 79%

Enhancing Scientific Figure Captioning Through Cross-modal Learning

Mateo Alejandro Rojas, Rafael Carranza

专题命中 图文多模态 :cross-modal(title);multimodal(abstract);分类 cs.CV

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16141 2024-06-25 cs.CV 79%

Multimodal Multilabel Classification by CLIP

Yanming Guo

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14830 2024-06-24 cs.CV 79%

CLIP-Decoder : ZeroShot Multilabel Classification using Multimodal CLIP Aligned Representation

Muhammad Ali, Salman Khan

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Accepted at ICCVW- VLAR

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14481 2024-06-21 cs.LG cs.AI cs.NE q-bio.NC 79%

Revealing Vision-Language Integration in the Brain with Multimodal Networks

Vighnesh Subramaniam, Colin Conwell, Christopher Wang, Gabriel Kreiman, Boris Katz, Ignacio Cases, Andrei Barbu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments ICML 2024; 23 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11820 2024-06-18 cs.CV 79%

Composing Object Relations and Attributes for Image-Text Matching

Khoi Pham, Chuong Huynh, Ser-Nam Lim, Abhinav Shrivastava

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted to CVPR'24

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04675 2024-06-10 cs.CV 79%

OVMR: Open-Vocabulary Recognition with Multi-Modal References

Zehong Ma, Shiliang Zhang, Longhui Wei, Qi Tian

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01906 2024-06-05 cs.CV cs.IR 79%

ProGEO: Generating Prompts through Image-Text Contrastive Learning for Visual Geo-localization

Chen Mao, Jingqi Hu

专题命中 图文多模态 :image-text(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20735 2024-06-03 cs.CV 79%

Language Augmentation in CLIP for Improved Anatomy Detection on Multi-modal Medical Images

Mansi Kakkar, Dattesh Shanbhag, Chandan Aladahalli, Gurunath Reddy M

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments $©$ 2024 IEEE. Accepted in 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02546 2024-05-30 cs.CV 79%

Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context Learning

Zhuo Huang, Chang Liu, Yinpeng Dong, Hang Su, Shibao Zheng, Tongliang Liu

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16301 2024-05-28 cs.CV cs.LG 79%

Active Learning for Finely-Categorized Image-Text Retrieval by Selecting Hard Negative Unpaired Samples

Dae Ung Jo, Kyuewang Lee, JaeHo Chung, Jin Young Choi

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12456 2024-05-22 eess.IV cs.CV cs.LG 79%

Mutual Information Analysis in Multimodal Learning Systems

Hadi Hadizadeh, S. Faegheh Yeganli, Bahador Rashidi, Ivan V. Bajić

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 6 pages, 7 figures, IEEE MIPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00384 2024-05-22 cs.CV 79%

TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias

Sanghyun Jo, Soohyun Ryu, Sungyub Kim, Eunho Yang, Kyungsu Kim

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11496 2024-05-21 cs.CV cs.IR 79%

DEMO: A Statistical Perspective for Efficient Image-Text Matching

Fan Zhang, Xian-Sheng Hua, Chong Chen, Xiao Luo

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09981 2024-05-17 cs.CV 79%

Adversarial Robustness for Visual Grounding of Multimodal Large Language Models

Kuofeng Gao, Yang Bai, Jiawang Bai, Yong Yang, Shu-Tao Xia

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments ICLR 2024 Workshop on Reliable and Responsible Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05949 2024-05-10 cs.CV 79%

CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts

Jiachen Li, Xinyao Wang, Sijie Zhu, Chia-Wen Kuo, Lu Xu, Fan Chen, Jitesh Jain, Humphrey Shi, Longyin Wen

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00962 2024-05-03 cs.CV 79%

FITA: Fine-grained Image-Text Aligner for Radiology Report Generation

Honglong Yang, Hui Tang, Xiaomeng Li

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments 11 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00029 2024-05-02 cs.CV cs.IR 79%

Automatic Creative Selection with Cross-Modal Matching

Alex Kim, Jia Huang, Rob Monarch, Jerry Kwac, Anikesh Kamath, Parmeshwar Khurd, Kailash Thiyagarajan, Goodman Gu

专题命中 图文多模态 :cross-modal(title);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15655 2024-04-25 cs.CV 79%

Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering

Jiawei Yao, Qi Qian, Juhua Hu

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2024. Project page: https://github.com/Alexander-Yao/Multi-MaP

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11864 2024-04-25 cs.CV 79%

Progressive Multi-modal Conditional Prompt Tuning

Xiaoyu Qiu, Hao Feng, Yuechen Wang, Wengang Zhou, Houqiang Li

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09797 2024-04-16 cs.CV 79%

TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Bozhi Luan, Hao Feng, Hong Chen, Yonghui Wang, Wengang Zhou, Houqiang Li

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14991 2024-04-15 cs.CV 79%

FoodLMM: A Versatile Food Assistant using Large Multi-modal Model

Yuehao Yin, Huiyan Qi, Bin Zhu, Jingjing Chen, Yu-Gang Jiang, Chong-Wah Ngo

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04676 2024-04-10 cs.CV eess.IV 79%

Stacked Cross-modal Feature Consolidation Attention Networks for Image Captioning

Mozhgan Pourkeshavarz, Shahabedin Nabavi, Mohsen Ebrahimi Moghaddam, Mehrnoush Shamsfard

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Journal ref Multimedia Tools and Applications, Volume 83, pages 12209-12233, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04960 2024-04-09 cs.CV 79%

PairAug: What Can Augmented Image-Text Pairs Do for Radiology?

Yutong Xie, Qi Chen, Sinuo Wang, Minh-Son To, Iris Lee, Ee Win Khoo, Kerolos Hendy, Daniel Koh, Yong Xia, Qi Wu

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted to CVPR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.07633 2024-04-09 cs.CL cs.LG 79%

Interpretable Detection of Out-of-Context Misinformation with Neural-Symbolic-Enhanced Large Multimodal Model

Yizhou Zhang, Loc Trinh, Defu Cao, Zijun Cui, Yan Liu

专题命中 图文多模态 :multimodal(title);cross-modal(abstract);分类 cs.CL

Comments 9 Pages, 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04231 2024-04-08 cs.CV 79%

Image-Text Co-Decomposition for Text-Supervised Semantic Segmentation

Ji-Jia Wu, Andy Chia-Hao Chang, Chieh-Yu Chuang, Chun-Pei Chen, Yu-Lun Liu, Min-Hung Chen, Hou-Ning Hu, Yung-Yu Chuang, Yen-Yu Lin

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01779 2024-04-01 cs.CV 79%

HallE-Control: Controlling Object Hallucination in Large Multimodal Models

Bohan Zhai, Shijia Yang, Chenfeng Xu, Sheng Shen, Kurt Keutzer, Chunyuan Li, Manling Li

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Our code is publicly available at https://github.com/bronyayang/HallE_Control

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02918 2024-03-21 cs.CV 79%

Multimodal Prompt Perceiver: Empower Adaptiveness, Generalizability and Fidelity for All-in-One Image Restoration

Yuang Ai, Huaibo Huang, Xiaoqiang Zhou, Jiexiang Wang, Ran He

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 13 pages, 8 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11593 2024-03-19 cs.CV 79%

End-to-end multi-modal product matching in fashion e-commerce

Sándor Tóth, Stephen Wilson, Alexia Tsoukara, Enric Moreu, Anton Masalovich, Lars Roemheld

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 9 pages, submitted to SIGKDD

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12596 2024-03-18 cs.CV 79%

UniHDA: A Unified and Versatile Framework for Multi-Modal Hybrid Domain Adaptation

Hengjia Li, Yang Liu, Yuqi Lin, Zhanwei Zhang, Yibo Zhao, weihang Pan, Tu Zheng, Zheng Yang, Yuchun Jiang, Boxi Wu, Deng Cai

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09027 2024-03-15 cs.CV 79%

VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework

Chris Kelly, Luhui Hu, Bang Yang, Yu Tian, Deshun Yang, Cindy Yang, Zaoshan Huang, Zihao Li, Jiayin Hu, Yuexian Zou

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 17 pages, 5 figures, and 1 table. arXiv admin note: substantial text overlap with arXiv:2311.10125

详情

展开后加载摘要…

URL PDF HTML 收藏