arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46294 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4676 篇

2310.20343 2023-11-06 cs.IR cs.MM 79%

Large Multi-modal Encoders for Recommendation

Zixuan Yi, Zijun Long, Iadh Ounis, Craig Macdonald, Richard Mccreadie

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.13959 2023-10-27 cs.CV 79%

Dynamic MDETR: A Dynamic Multimodal Transformer Decoder for Visual Grounding

Fengyuan Shi, Ruopeng Gao, Weilin Huang, Limin Wang

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) in October 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13398 2023-10-23 cs.CV 79%

OpenAnnotate3D: Open-Vocabulary Auto-Labeling System for Multi-modal 3D Data

Yijie Zhou, Likun Cai, Xianhui Cheng, Zhongxue Gan, Xiangyang Xue, Wenchao Ding

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments The source code will be released at https://github.com/Fudan-ProjectTitan/OpenAnnotate3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14539 2023-10-12 cs.CR cs.CL 79%

Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models

Erfan Shayegani, Yue Dong, Nael Abu-Ghazaleh

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05109 2023-10-10 cs.CV 79%

Lightweight In-Context Tuning for Multimodal Unified Models

Yixin Chen, Shuai Zhang, Boran Han, Jiaya Jia

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01936 2023-10-04 cs.CV 79%

Constructing Image-Text Pair Dataset from Books

Yamato Okamoto, Haruto Toyonaga, Yoshihisa Ijiri, Hirokatsu Kataoka

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2023 workshop, Towards the Next Generation of Computer Vision Datasets: General DataCentric Submission Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00108 2023-10-03 cs.LG cs.CV 79%

Practical Membership Inference Attacks Against Large-Scale Multi-Modal Models: A Pilot Study

Myeongseob Ko, Ming Jin, Chenguang Wang, Ruoxi Jia

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments International Conference on Computer Vision (ICCV) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18009 2023-09-26 cs.CV 79%

Multi-Modal Face Stylization with a Generative Prior

Mengtian Li, Yi Dong, Minxuan Lin, Haibin Huang, Pengfei Wan, Chongyang Ma

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12110 2023-09-22 cs.CV 79%

Exploiting CLIP-based Multi-modal Approach for Artwork Classification and Retrieval

Alberto Baldrati, Marco Bertini, Tiberio Uricchio, Alberto Del Bimbo

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments Proc. of Florence Heri-Tech 2022: The Future of Heritage Science and Technologies: ICT and Digital Heritage, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10206 2023-09-20 cs.CV 79%

Image-Text Pre-Training for Logo Recognition

Mark Hubenthal, Suren Kumar

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments 8 pages, 5 figures, 4 tables

Journal ref Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2023, pp. 1145-1154

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04790 2023-09-12 cs.CL 79%

MMHQA-ICL: Multimodal In-context Learning for Hybrid Question Answering over Text, Tables and Images

Weihao Liu, Fangyu Lei, Tongxu Luo, Jiahe Lei, Shizhu He, Jun Zhao, Kang Liu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.18232 2023-08-16 cs.CV 79%

DIME-FM: DIstilling Multimodal and Efficient Foundation Models

Ximeng Sun, Pengchuan Zhang, Peizhao Zhang, Hardik Shah, Kate Saenko, Xide Xia

专题命中 图文多模态 :multimodal(title);image-text(abstract);分类 cs.CV

Comments Accepted to ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.00847 2023-08-15 cs.CV 79%

MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the Wild

Yuanyuan Liu, Wei Dai, Chuanxu Feng, Wenbin Wang, Guanghao Yin, Jiabei Zeng, Shiguang Shan

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments This paper has been accepted by ACM MM'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06262 2023-08-14 cs.LG cs.AI 79%

Foundation Model is Efficient Multimodal Multitask Model Selector

Fanqing Meng, Wenqi Shao, Zhanglin Peng, Chonghe Jiang, Kaipeng Zhang, Yu Qiao, Ping Luo

专题命中 图文多模态 :multimodal(title);multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12964 2023-08-07 cs.CV 79%

Text-based Person Search without Parallel Image-Text Data

Yang Bai, Jingyao Wang, Min Cao, Chen Chen, Ziqiang Cao, Liqiang Nie, Min Zhang

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.05920 2023-07-13 eess.IV cs.CV cs.LG 79%

Unified Medical Image-Text-Label Contrastive Learning With Continuous Prompt

Yuhao Wang

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.02677 2023-07-07 cs.CV 79%

Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Teng Wang, Jinrui Zhang, Junjie Fei, Hao Zheng, Yunlong Tang, Zhe Li, Mingqi Gao, Shanshan Zhao

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Tech-report

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07198 2023-06-13 cs.CL 79%

A Survey of Vision-Language Pre-training from the Lens of Multimodal Machine Translation

Jeremy Gwinnup, Kevin Duh

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.02348 2023-06-06 cs.CL 79%

Leverage Points in Modality Shifts: Comparing Language-only and Multimodal Word Representations

Aleksey Tikhonov, Lisa Bylinina, Denis Paperno

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted for StarSEM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12256 2023-05-26 cs.CL 79%

Scene Graph as Pivoting: Inference-time Image-free Unsupervised Multimodal Machine Translation with Visual Scene Hallucination

Hao Fei, Qian Liu, Meishan Zhang, Min Zhang, Tat-Seng Chua

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.04524 2023-05-09 cs.CV 79%

Scene Text Recognition with Image-Text Matching-guided Dictionary

Jiajun Wei, Hongjian Zhan, Xiao Tu, Yue Lu, Umapada Pal

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted at ICDAR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05173 2023-04-12 cs.CV cs.LG 79%

Improving Image Recognition by Retrieving from Web-Scale Image-Text Data

Ahmet Iscen, Alireza Fathi, Cordelia Schmid

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00785 2023-03-28 cs.CV 79%

Learning to Generate Text-grounded Mask for Open-world Semantic Segmentation from Only Image-Text Pairs

Junbum Cha, Jonghwan Mun, Byungseok Roh

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments CVPR 2023 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13340 2023-03-24 cs.LG cs.CV 79%

Increasing Textual Context Size Boosts Medical Image-Text Matching

Idan Glassberg, Tom Hope

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12997 2023-03-24 cs.CV 79%

FER-former: Multi-modal Transformer for Facial Expression Recognition

Yande Li, Mingjie Wang, Minglun Gong, Yonggang Lu, Li Liu

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07122 2023-03-14 cs.CV 79%

ContextCLIP: Contextual Alignment of Image-Text pairs on CLIP visual representations

Chanda Grover, Indra Deep Mastan, Debayan Gupta

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments 11 Pages, 7 Figures, 2 Tables, ICVGIP

Journal ref ICVGIP, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07254 2023-03-03 cs.CV cs.LG 79%

The Role of Local Alignment and Uniformity in Image-Text Contrastive Learning on Medical Images

Philip Müller, Georgios Kaissis, Daniel Rueckert

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments NeurIPS 2022 Workshop: Self-Supervised Learning - Theory and Practice (Reason for updated version: correction of a typo in Eq. (2) and (3))

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02908 2023-02-07 cs.CV cs.IR 79%

LexLIP: Lexicon-Bottlenecked Language-Image Pre-Training for Large-Scale Image-Text Retrieval

Ziyang luo, Pu Zhao, Can Xu, Xiubo Geng, Tao Shen, Chongyang Tao, Jing Ma, Qingwen lin, Daxin Jiang

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.05453 2023-02-07 cs.CL 79%

It's Just a Matter of Time: Detecting Depression with Time-Enriched Multimodal Transformers

Ana-Maria Bucur, Adrian Cosma, Paolo Rosso, Liviu P. Dinu

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at ECIR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06844 2023-01-18 cs.CV 79%

USER: Unified Semantic Enhancement with Momentum Contrast for Image-Text Retrieval

Yan Zhang, Zhong Ji, Di Wang, Yanwei Pang, Xuelong Li

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏