arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46352 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4682 篇

2305.18171 2024-04-10 cs.CV cs.LG 83%

Improved Probabilistic Image-Text Representations

Sanghyuk Chun

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

Comments ICLR 2024 camera-ready; Code: https://github.com/naver-ai/pcmepp. Project page: https://naver-ai.github.io/pcmepp/. 30 pages, 2.2 MB

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08498 2024-04-09 cs.CV 83%

Extending CLIP's Image-Text Alignment to Referring Image Segmentation

Seoyeon Kim, Minguk Kang, Dongwon Kim, Jaesik Park, Suha Kwak

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

Comments NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01700 2024-04-04 cs.CV 83%

MotionChain: Conversational Motion Controllers via Multimodal Prompts

Biao Jiang, Xin Chen, Chi Zhang, Fukun Yin, Zhuoyuan Li, Gang YU, Jiayuan Fan

专题命中 图文多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 14 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.10185 2024-03-27 cs.CV 83%

ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights

Weixian Lei, Yixiao Ge, Jianfeng Zhang, Dylan Sun, Kun Yi, Ying Shan, Mike Zheng Shou

专题命中 图文多模态 :omni-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments 19 pages, 4 figures and 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10887 2024-03-19 cs.CV 83%

LuoJiaHOG: A Hierarchy Oriented Geo-aware Image Caption Dataset for Remote Sensing Image-Text Retrival

Yuanxin Zhao, Mi Zhang, Bingnan Yang, Zhan Zhang, Jiaju Kang, Jianya Gong

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02110 2024-03-12 cs.CV 83%

Sieve: Multimodal Dataset Pruning Using Image Captioning Models

Anas Mahmoud, Mostafa Elhoushi, Amro Abbas, Yu Yang, Newsha Ardalani, Hugh Leather, Ari Morcos

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted in CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04547 2024-03-08 cs.LG cs.AI 83%

CLIP the Bias: How Useful is Balancing Data in Multimodal Learning?

Ibrahim Alabdulmohsin, Xiao Wang, Andreas Steiner, Priya Goyal, Alexander D'Amour, Xiaohua Zhai

专题命中 图文多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments 32 pages, 20 figures, 7 tables

Journal ref ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03003 2024-03-06 cs.CV 83%

Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Gen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng, Xiaoshuai Sun, Rongrong Ji

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15654 2024-02-27 cs.CL 83%

Exploring Failure Cases in Multimodal Reasoning About Physical Dynamics

Sadaf Ghaffari, Nikhil Krishnaswamy

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

Comments 10 pages, 10 figures, Proceedings of AAAI Spring Symposium: Empowering Machine Learning and Large Language Models with Domain and Commonsense Knowledge (MAKE). AAAI (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08360 2024-02-14 cs.CV 83%

Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks

Jusung Lee, Sungguk Cha, Younghyun Lee, Cheoljong Yang

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.04655 2024-02-07 cs.CR cs.CV 83%

VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models

Ziyi Yin, Muchao Ye, Tianrong Zhang, Tianyu Du, Jinguo Zhu, Han Liu, Jinghui Chen, Ting Wang, Fenglong Ma

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2023, 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09172 2024-01-19 cs.CV cs.LG 83%

Hyperbolic Image-Text Representations

Karan Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson, Ramakrishna Vedantam

专题命中 图文多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV

Comments ICML 2023 (v3: Add link to code in abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09725 2024-01-19 cs.IR cs.MM 83%

Enhancing Image-Text Matching with Adaptive Feature Aggregation

Zuhui Wang, Yunting Yin, I. V. Ramakrishnan

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.MM

Comments Accepted by ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17468 2024-01-09 cs.CV cs.LG 83%

Cross-modal Active Complementary Learning with Self-refining Correspondence

Yang Qin, Yuan Sun, Dezhong Peng, Joey Tianyi Zhou, Xi Peng, Peng Hu

专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments This paper is accepted by NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01736 2024-01-05 cs.CV 83%

Few-shot Adaptation of Multi-modal Foundation Models: A Survey

Fan Liu, Tianshu Zhang, Wenwen Dai, Wenwen Cai, Xiaocong Zhou, Delong Chen

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14233 2023-12-25 cs.CV 83%

VCoder: Versatile Vision Encoders for Multimodal Large Language Models

Jitesh Jain, Jianwei Yang, Humphrey Shi

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Project Page: https://praeclarumjj3.github.io/vcoder/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08084 2023-12-18 cs.AI 83%

A Novel Energy based Model Mechanism for Multi-modal Aspect-Based Sentiment Analysis

Tianshuo Peng, Zuchao Li, Ping Wang, Lefei Zhang, Hai Zhao

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.AI

Comments AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.04364 2023-12-12 cs.MM 83%

ITportrait: Image-Text Coupled 3D Portrait Domain Adaptation

Xiangwen Deng, Yufeng Wang, Yuanhao Cai, Jingxiang Sun, Yebin Liu, Haoqian Wang

专题命中 图文多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05863 2023-11-13 cs.CR cs.CV 83%

Watermarking Vision-Language Pre-trained Models for Multi-modal Embedding as a Service

Yuanmin Tang, Jing Yu, Keke Gai, Xiangyan Qu, Yue Hu, Gang Xiong, Qi Wu

专题命中 图文多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05437 2023-11-10 cs.CV cs.AI cs.CL cs.LG cs.MM 83%

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang, Jianfeng Gao, Chunyuan Li

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments 25 pages, 25M file size. Project Page: https://llava-vl.github.io/llava-plus/

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10350 2023-10-27 cs.LG cs.CV 83%

Improving Multimodal Datasets with Image Captioning

Thao Nguyen, Samir Yitzhak Gadre, Gabriel Ilharco, Sewoong Oh, Ludwig Schmidt

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments Accepted at NeurIPS 2023 Datasets & Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08368 2023-10-13 cs.CV 83%

Mapping Memes to Words for Multimodal Hateful Meme Classification

Giovanni Burbi, Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini, Alberto Del Bimbo

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

Comments ICCV2023 CLVL Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.01615 2023-09-19 cs.CV 83%

ConTEXTual Net: A Multimodal Vision-Language Model for Segmentation of Pneumothorax

Zachary Huemann, Xin Tie, Junjie Hu, Tyler J. Bradshaw

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03921 2023-09-11 cs.CV 83%

C-CLIP: Contrastive Image-Text Encoders to Close the Descriptive-Commentative Gap

William Theisen, Walter Scheirer

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV

Comments 11 Pages, 5 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10313 2023-09-06 cs.CL 83%

Beyond Triplet: Leveraging the Most Data for Multimodal Machine Translation

Yaoming Zhu, Zewei Sun, Shanbo Cheng, Luyang Huang, Liwei Wu, Mingxuan Wang

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

Comments 8 pages, ACL 2023 Finding

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16527 2023-08-22 cs.IR cs.CV 83%

OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander M. Rush, Douwe Kiela, Matthieu Cord, Victor Sanh

专题命中 图文多模态 :image-text(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.13181 2023-08-15 cs.LG cs.CV 83%

Sample-Specific Debiasing for Better Image-Text Models

Peiqi Wang, Yingcheng Liu, Ching-Yun Ko, William M. Wells, Seth Berkowitz, Steven Horng, Polina Golland

专题命中 图文多模态 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Machine Learning for Healthcare Conference 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04832 2023-07-07 cs.MM cs.SI 83%

Unifying Multimodal Source and Propagation Graph for Rumour Detection on Social Media with Missing Features

Tsun-Hin Cheung, Kin-Man Lam

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.15195 2023-07-04 cs.CV 83%

Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, Rui Zhao

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.01758 2023-05-26 cs.CV 83%

Improving Zero-shot Generalization and Robustness of Multi-modal Models

Yunhao Ge, Jie Ren, Andrew Gallagher, Yuxiao Wang, Ming-Hsuan Yang, Hartwig Adam, Laurent Itti, Balaji Lakshminarayanan, Jiaping Zhao

专题命中 图文多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV

Comments CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏