arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46122 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4657 篇

2112.00599 2021-12-02 cs.CV cs.LG 57%

An implementation of the "Guess who?" game using CLIP

Arnau Martí Sarri, Victor Rodriguez-Fernandez

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Code available at https://github.com/ArnauDIMAI/CLIP-GuessWho

Journal ref Intelligent Data Engineering and Automated Learning (IDEAL 2021). Lecture Notes in Computer Science, vol 13113

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00331 2021-12-02 cs.MM 57%

Mutltimodal AI Companion for Interactive Fairytale Co-creation

Ruiyang Liu, Predrag K. Nikolic

专题命中 图文多模态 :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.14745 2021-11-30 cs.CV 57%

A Simple Long-Tailed Recognition Baseline via Vision-Language Model

Teli Ma, Shijie Geng, Mengmeng Wang, Jing Shao, Jiasen Lu, Hongsheng Li, Peng Gao, Yu Qiao

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.13118 2021-10-26 cs.CV 57%

TRIE: End-to-End Text Reading and Information Extraction for Document Understanding

Peng Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Jing Lu, Liang Qiao, Yi Niu, Fei Wu

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to ACM MM2020. Code is available at https://davar-lab.github.io/publication.html or https://github.com/hikopensource/DAVAR-Lab-OCR

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.03743 2021-09-15 cs.CV 57%

Visual News: Benchmark and Challenges in News Image Captioning

Fuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente Ordonez

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments 9 pages, 5 figures, accepted to EMNLP2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.02955 2021-09-08 cs.CV 57%

Sensor-Augmented Egocentric-Video Captioning with Dynamic Modal Attention

Katsuyuki Nakamura, Hiroki Ohashi, Mitsuhiro Okada

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ACM Multimedia (ACMMM) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.06946 2021-08-11 cs.CV 57%

MiniVLM: A Smaller and Faster Vision-Language Model

Jianfeng Wang, Xiaowei Hu, Pengchuan Zhang, Xiujun Li, Lijuan Wang, Lei Zhang, Jianfeng Gao, Zicheng Liu

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.14774 2021-06-01 cs.CL 57%

LIIR at SemEval-2021 task 6: Detection of Persuasion Techniques In Texts and Images using CLIP features

Erfan Ghadery, Damien Sileo, Marie-Francine Moens

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.09906 2021-05-25 cs.CV 57%

Empirical Analysis of Image Caption Generation using Deep Learning

Aditya Bhattacharya, Eshwar Shamanna Girishekar, Padmakar Anil Deshpande

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Withdrawing for further updates to the work

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08945 2021-04-20 cs.CV 57%

Data-Efficient Language-Supervised Zero-Shot Learning with Self-Distillation

Ruizhe Cheng, Bichen Wu, Peizhao Zhang, Peter Vajda, Joseph E. Gonzalez

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments 4 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.16110 2021-04-16 cs.CV 57%

Kaleido-BERT: Vision-Language Pre-training on Fashion Domain

Mingchen Zhuge, Dehong Gao, Deng-Ping Fan, Linbo Jin, Ben Chen, Haoming Zhou, Minghui Qiu, Ling Shao

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments CVPR2021 Accepted. Code: https://github.com/mczhuge/Kaleido-BERT

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.07192 2021-02-16 cs.CV 57%

Improved Bengali Image Captioning via deep convolutional neural network based encoder-decoder model

Mohammad Faiyaz Khan, S. M. Sadiq-Ur-Rahman Shifath, Md. Saiful Islam

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted in "IJCACI 2020: International Joint Conference on Advances in Computational Intelligence"

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.03942 2021-02-09 cs.CV 57%

Iconographic Image Captioning for Artworks

Eva Cetinic

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments Accepted at Workshop on Fine Art Pattern Extraction and Recognition (FAPER), ICPR, 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.08000 2021-01-21 cs.CV 57%

Macroscopic Control of Text Generation for Image Captioning

Zhangzi Zhu, Tianlei Wang, Hong Qu

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.13192 2020-12-24 cs.CV 57%

TIME: Text and Image Mutual-Translation Adversarial Networks

Bingchen Liu, Kunpeng Song, Yizhe Zhu, Gerard de Melo, Ahmed Elgammal

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments AAAI-2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.07061 2020-12-15 cs.CV 57%

Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer Network

Jiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen, Gen Luo, Yongjian Wu, Yue Gao, Rongrong Ji

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted at AAAI 2021 (preprint version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.04638 2020-12-09 cs.CV 57%

TAP: Text-Aware Pre-training for Text-VQA and Text-Caption

Zhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin, Dinei Florencio, Lijuan Wang, Cha Zhang, Lei Zhang, Jiebo Luo

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.10966 2020-08-26 cs.CV 57%

In-Home Daily-Life Captioning Using Radio Signals

Lijie Fan, Tianhong Li, Yuan Yuan, Dina Katabi

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments ECCV 2020. The first two authors contributed equally to this paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.11693 2020-08-13 cs.CV 57%

Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020

Teng Wang, Huicheng Zheng, Mingjing Yu

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments Second-place solution to TASK 2 (Dense video captioning) in ActivityNet Challenge 2020. Code is available at https://github.com/ttengwang/dense-video-captioning-pytorch

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.03098 2020-07-21 cs.CV 57%

Connecting Vision and Language with Localized Narratives

Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut, Vittorio Ferrari

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments ECCV 2020 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.07883 2020-04-02 cs.CV 57%

Vision-Language Navigation with Self-Supervised Auxiliary Reasoning Tasks

Fengda Zhu, Yi Zhu, Xiaojun Chang, Xiaodan Liang

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.14080 2020-04-01 cs.CV 57%

X-Linear Attention Networks for Image Captioning

Yingwei Pan, Ting Yao, Yehao Li, Tao Mei

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments CVPR 2020; The source code and model are publicly available at: https://github.com/Panda-Peter/image-captioning

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.09536 2020-02-25 cs.CV 57%

Image to Language Understanding: Captioning approach

Madhavan Seshadri, Malavika Srikanth, Mikhail Belov

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.08960 2019-12-20 cs.CL 57%

Going Beneath the Surface: Evaluating Image Captioning for Grammaticality, Truthfulness and Diversity

Huiyuan Xie, Tom Sherborne, Alexander Kuhnle, Ann Copestake

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.08360 2019-12-19 cs.CL 57%

DMRM: A Dual-channel Multi-hop Reasoning Model for Visual Dialog

Feilong Chen, Fandong Meng, Jiaming Xu, Peng Li, Bo Xu, Jie Zhou

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted at AAAI 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.03083 2019-12-09 cs.CV 57%

Visual-Textual Association with Hardest and Semi-Hard Negative Pairs Mining for Person Search

Jing Ge, Guangyu Gao, Zhen Liu

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.09796 2019-12-09 cs.CV 57%

A Novel Technique for Evidence based Conditional Inference in Deep Neural Networks via Latent Feature Perturbation

Dinesh Khandelwal, Suyash Agrawal, Parag Singla, Chetan Arora

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.11059 2019-12-05 cs.CV 57%

Unified Vision-Language Pre-Training for Image Captioning and VQA

Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, Jianfeng Gao

专题命中 图文多模态 :image-text(abstract);分类 cs.CV

Comments AAAI 2020 camera-ready version. The code and the pre-trained models are available at https://github.com/LuoweiZhou/VLP

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.06475 2019-10-16 cs.CV 57%

Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive Style

Hongwei Ge, Zehang Yan, Kai Zhang, Mingde Zhao, Liang Sun

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

Comments ICCV 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.03609 2019-07-09 cs.CV 57%

Variational Context: Exploiting Visual and Textual Context for Grounding Referring Expressions

Yulei Niu, Hanwang Zhang, Zhiwu Lu, Shih-Fu Chang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted as regular paper in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI). Substantial text overlap with arXiv:1712.01892

详情

展开后加载摘要…

URL PDF HTML 收藏