arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3493 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3493 篇

2104.01894 2021-06-16 cs.CL cs.CV cs.IR cs.LG 62%

Talk, Don't Write: A Study of Direct Speech-Based Image Retrieval

Ramon Sanabria, Austin Waters, Jason Baldridge

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted to INTERSPEECH 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.01832 2021-04-06 cs.CV cs.AI 62%

Task-Independent Knowledge Makes for Transferable Representations for Generalized Zero-Shot Learning

Chaoqun Wang, Xuejin Chen, Shaobo Min, Xiaoyan Sun, Houqiang Li

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at AAAI2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.08673 2021-04-01 cs.CV cs.CL 62%

A Closer Look at the Robustness of Vision-and-Language Pre-trained Models

Linjie Li, Zhe Gan, Jingjing Liu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.04863 2021-03-09 cs.CV cs.AI 62%

From Hand-Perspective Visual Information to Grasp Type Probabilities: Deep Learning via Ranking Labels

Mo Han, Sezen Ya{ğ}mur Günay, İlkay Yıldız, Paolo Bonato, Cagdas D. Onal, Taşkın Padır, Gunar Schirner, Deniz Erdo{ğ}muş

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.09375 2021-02-19 cs.CV cs.IR cs.MM 62%

Hierarchical Similarity Learning for Language-based Product Image Retrieval

Zhe Ma, Fenghao Liu, Jianfeng Dong, Xiaoye Qu, Yuan He, Shouling Ji

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted by ICASSP 2021. Code and data will be available at https://github.com/liufh1/hsl

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.12350 2021-01-27 cs.CL cs.AI 62%

ActionBert: Leveraging User Actions for Semantic Understanding of User Interfaces

Zecheng He, Srinivas Sunkara, Xiaoxue Zang, Ying Xu, Lijuan Liu, Nevan Wichers, Gabriel Schubiner, Ruby Lee, Jindong Chen, Blaise Agüera y Arcas

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted to AAAI Conference on Artificial Intelligence (AAAI-21)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.08148 2020-12-16 cs.CL cs.AI 62%

A Response Retrieval Approach for Dialogue Using a Multi-Attentive Transformer

Matteo A. Senese, Alberto Benincasa, Barbara Caputo, Giuseppe Rizzo

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.07000 2020-12-15 cs.AI cs.CL 62%

KVL-BERT: Knowledge Enhanced Visual-and-Linguistic BERT for Visual Commonsense Reasoning

Dandan Song, Siyi Ma, Zhanchen Sun, Sicheng Yang, Lejian Liao

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.05107 2020-12-10 cs.CL cs.CV cs.LG 62%

Towards Zero-shot Cross-lingual Image Retrieval

Pranav Aggarwal, Ajinkya Kale

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.02467 2020-10-07 cs.CV cs.CL 62%

Learning Visual-Semantic Embeddings for Reporting Abnormal Findings on Chest X-rays

Jianmo Ni, Chun-Nan Hsu, Amilcare Gentili, Julian McAuley

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments 7 pages, 2 figures, to be published in Findings of EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.12212 2020-09-24 cs.CV cs.CL cs.IR cs.LG 62%

ZSCRGAN: A GAN-based Expectation Maximization Model for Zero-Shot Retrieval of Images from Textual Descriptions

Anurag Roy, Vinay Kumar Verma, Kripabandhu Ghosh, Saptarshi Ghosh

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted in CIKM-2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.00392 2020-03-03 cs.CV cs.AI 62%

Fine-grained Video-Text Retrieval with Hierarchical Graph Reasoning

Shizhe Chen, Yida Zhao, Qin Jin, Qi Wu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments To be appeared in CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.09461 2020-02-24 cs.CV cs.MM 62%

Fine-Grained Instance-Level Sketch-Based Video Retrieval

Peng Xu, Kun Liu, Tao Xiang, Timothy M. Hospedales, Zhanyu Ma, Jun Guo, Yi-Zhe Song

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.03712 2020-01-14 cs.CV cs.CL cs.LG 62%

MHSAN: Multi-Head Self-Attention Network for Visual Semantic Embedding

Geondo Park, Chihye Han, Wonjun Yoon, Daeshik Kim

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted by the 2020 IEEE Winter Conference on Applications of Computer Vision (WACV 20), 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.10531 2019-11-26 cs.CV cs.MM eess.IV 62%

A Proposal-based Approach for Activity Image-to-Video Retrieval

Ruicong Xu, Li Niu, Jianfu Zhang, Liqing Zhang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments The Thirty-Fourth AAAI Conference on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.10097 2019-11-25 cs.LG cs.CL cs.CV 62%

HAL: Improved Text-Image Matching by Mitigating Visual Semantic Hubs

Fangyu Liu, Rongtian Ye, Xun Wang, Shuaipeng Li

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments AAAI-20 (to appear)

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.05978 2019-11-15 cs.CV cs.CL cs.LG 62%

HUSE: Hierarchical Universal Semantic Embeddings

Pradyumna Narayana, Aniket Pednekar, Abishek Krishnamoorthy, Kazoo Sone, Sugato Basu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.11119 2019-10-25 cs.CV cs.CL cs.LG 62%

Designovel's system description for Fashion-IQ challenge 2019

Jianri Li, Jae-whan Lee, Woo-sang Song, Ki-young Shin, Byung-hyun Go

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.01205 2019-06-05 cs.LG cs.CL cs.CV 62%

A Strong and Robust Baseline for Text-Image Matching

Fangyu Liu, Rongtian Ye

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments 6 pages (excluding references); 2019 ACL Student Research Workshop (to appear)

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.05521 2019-04-30 cs.CV cs.CL cs.LG 62%

UniVSE: Robust Visual Semantic Embeddings via Structured Semantic Representations

Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang, Lei Li, Weiwei Sun, Wei-Ying Ma

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments v1 is the full version which is accepted by CVPR 2019. v2 is the short version accepted by NAACL 2019 SpLU-RoboNLP workshop (in non-archival proceedings)

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.09362 2018-11-27 cs.CL cs.AI 62%

Words Can Shift: Dynamically Adjusting Word Representations Using Nonverbal Behaviors

Yansen Wang, Ying Shen, Zhun Liu, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted by AAAI2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.10348 2018-06-28 cs.CL cs.CV 62%

Learning Visually-Grounded Semantics from Contrastive Adversarial Samples

Haoyue Shi, Jiayuan Mao, Tete Xiao, Yuning Jiang, Jian Sun

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL

Comments To Appear at COLING 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.08495 2018-03-23 cs.CV cs.AI cs.GR cs.LG 62%

Text2Shape: Generating Shapes from Natural Language by Learning Joint Embeddings

Kevin Chen, Christopher B. Choy, Manolis Savva, Angel X. Chang, Thomas Funkhouser, Silvio Savarese

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.07199 2017-12-21 cs.DB cs.AI cs.CL cs.NE 62%

Cognitive Database: A Step towards Endowing Relational Databases with Artificial Intelligence Capabilities

Rajesh Bordawekar, Bortik Bandyopadhyay, Oded Shmueli

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.05851 2017-08-22 cs.CV cs.IR cs.MM 62%

Image2song: Song Retrieval via Bridging Image Content and Lyric Words

Xuelong Li, Di Hu, Xiaoqiang Lu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.MM

Comments 13 pages, 13 figures, accepted by ICCV 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.02531 2017-08-09 cs.CV cs.AI 62%

Deep Binaries: Encoding Semantic-Rich Cues for Efficient Textual-Visual Cross Retrieval

Yuming Shen, Li Liu, Ling Shao, Jingkuan Song

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2017 as a conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.05908 2017-03-21 cs.CV cs.CL cs.LG 62%

Learning Robust Visual-Semantic Embeddings

Yao-Hung Hubert Tsai, Liang-Kang Huang, Ruslan Salakhutdinov

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1505.04870 2016-09-21 cs.CV cs.CL 62%

Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, Svetlana Lazebnik

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1603.08474 2016-03-29 cs.CL cs.CV cs.LG cs.NE 62%

Deep Embedding for Spatial Role Labeling

Oswaldo Ludwig, Xiao Liu, Parisa Kordjamshidi, Marie-Francine Moens

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.03478 2016-01-15 cs.LG cs.CL cs.CV 62%

Deep Learning Applied to Image and Text Matching

Afroze Ibrahim Baqapuri

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏