arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3463 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3463 篇

2301.04647 2023-06-21 cs.CV cs.CL 81%

EXIF as Language: Learning Cross-Modal Associations Between Images and Camera Metadata

Chenhao Zheng, Ayush Shrivastava, Andrew Owens

专题命中 跨模态检索 :cross-modal(title);multimodal(abstract);分类 cs.CV、cs.CL

Comments CVPR 2023 (Highlight). Project link: http://hellomuffin.github.io/exif-as-language

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.02549 2023-06-14 cs.CL cs.CV cs.LG 81%

FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information Extraction

Chen-Yu Lee, Chun-Liang Li, Hao Zhang, Timothy Dozat, Vincent Perot, Guolong Su, Xiang Zhang, Kihyuk Sohn, Nikolai Glushnev, Renshen Wang, Joshua Ainslie, Shangbang Long, Siyang Qin, Yasuhisa Fujii, Nan Hua, Tomas Pfister

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted to ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12561 2023-06-07 cs.CV cs.CL cs.LG 81%

Retrieval-Augmented Multimodal Language Modeling

Michihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Rich James, Jure Leskovec, Percy Liang, Mike Lewis, Luke Zettlemoyer, Wen-tau Yih

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Published at ICML 2023. Blog post available at https://cs.stanford.edu/~myasu/blog/racm3/

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01004 2023-06-05 cs.CL cs.AI 81%

AoM: Detecting Aspect-oriented Information for Multimodal Aspect-Based Sentiment Analysis

Ru Zhou, Wenya Guo, Xumeng Liu, Shenglong Yu, Ying Zhang, Xiaojie Yuan

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Findings of ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.11618 2023-04-25 cs.CL cs.AI 81%

Modality-Aware Negative Sampling for Multi-modal Knowledge Graph Embedding

Yichi Zhang, Mingyang Chen, Wen Zhang

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by IJCNN2023. Code is released in https://github.com/zjukg/MANS

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03717 2023-04-10 cs.LG cs.CL cs.CV 81%

On the Importance of Contrastive Loss in Multimodal Learning

Yunwei Ren, Yuanzhi Li

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00448 2023-04-04 cs.CV cs.MM 81%

The style transformer with common knowledge optimization for image-text retrieval

Wenrui Li, Zhengyu Ma, Jinqiao Shi, Xiaopeng Fan

专题命中 跨模态检索 :image-text(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07740 2023-03-15 cs.CV cs.CL 81%

Efficient Image-Text Retrieval via Keyword-Guided Pre-Screening

Min Cao, Yang Bai, Jingyao Wang, Ziqiang Cao, Liqiang Nie, Min Zhang

专题命中 跨模态检索 :image-text(title,abstract);分类 cs.CV、cs.CL

Comments 11 pages, 7 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.01781 2023-03-06 cs.CV cs.CL 81%

Meme Sentiment Analysis Enhanced with Multimodal Spatial Encoding and Facial Embedding

Muzhaffar Hazman, Susan McKeever, Josephine Griffith

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Published as chapter in ISBN:978-3-031-26438-2

Journal ref In: Longo, L., OReilly, R. (eds) Artificial Intelligence and Cognitive Science. AICS 2022. Communications in Computer and Information Science, vol 1662. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06468 2023-02-15 cs.AI cs.CL cs.LG 81%

Contrastive Multimodal Learning for Emergence of Graphical Sensory-Motor Communication

Tristan Karch, Yoann Lemesle, Romain Laroche, Clément Moulin-Frier, Pierre-Yves Oudeyer

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.04366 2023-01-12 cs.CL cs.IR cs.LG cs.MM 81%

Multimodal Inverse Cloze Task for Knowledge-based Visual Question Answering

Paul Lerner, Olivier Ferret, Camille Guinaudeau

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments Accepted at ECIR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.14322 2023-01-02 cs.IR cs.AI cs.MM 81%

BagFormer: Better Cross-Modal Retrieval via bag-wise interaction

Haowen Hou, Xiaopeng Yan, Yigeng Zhang, Fengzong Lian, Zhanhui Kang

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.AI、cs.MM

Comments 8 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11319 2022-10-21 cs.CV cs.MM 81%

Image-Text Retrieval with Binary and Continuous Label Supervision

Zheng Li, Caili Guo, Zerun Feng, Jenq-Neng Hwang, Ying Jin, Yufeng Zhang

专题命中 跨模态检索 :image-text(title,abstract);分类 cs.CV、cs.MM

Comments 13 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03162 2022-10-21 cs.CV cs.CL 81%

Embedding Arithmetic of Multimodal Queries for Image Retrieval

Guillaume Couairon, Matthieu Cord, Matthijs Douze, Holger Schwenk

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments accepted at O-DRUM (CVPR workshop 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05040 2022-10-06 cs.CL cs.AI 81%

SANCL: Multimodal Review Helpfulness Prediction with Selective Attention and Natural Contrastive Learning

Wei Han, Hui Chen, Zhen Hai, Soujanya Poria, Lidong Bing

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted as a long paper at COLING 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12599 2022-09-27 cs.CV cs.AI 81%

Deep Manifold Hashing: A Divide-and-Conquer Approach for Semi-Paired Unsupervised Cross-Modal Retrieval

Yufeng Shi, Xinge You, Jiamiao Xu, Feng Zheng, Qinmu Peng, Weihua Ou

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08012 2022-09-26 cs.CL cs.AI cs.LG 81%

CascadER: Cross-Modal Cascading for Knowledge Graph Link Prediction

Tara Safavi, Doug Downey, Tom Hope

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CL、cs.AI

Comments AKBC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.06515 2022-09-20 cs.CV cs.MM 81%

Learning to Evaluate Performance of Multi-modal Semantic Localization

Zhiqiang Yuan, Wenkai Zhang, Chongyang Li, Zhaoying Pan, Yongqiang Mao, Jialiang Chen, Shouke Li, Hongqi Wang, Xian Sun

专题命中 跨模态检索 :multi-modal(title);cross-modal(abstract);分类 cs.CV、cs.MM

Comments 19 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.07084 2022-09-16 cs.AI cs.CL 81%

Knowledge Graph Completion with Pre-trained Multimodal Transformer and Twins Negative Sampling

Yichi Zhang, Wen Zhang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by KDD 2022 Undergraduate Consortium

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.00682 2022-09-05 cs.CV cs.AI cs.GR cs.LG 81%

Zero-Shot Multi-Modal Artist-Controlled Retrieval and Exploration of 3D Object Sets

Kristofer Schlachter, Benjamin Ahlbrand, Zhu Wang, Valerio Ortenzi, Ken Perlin

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.14326 2022-08-31 cs.CV cs.AI 81%

GaitFi: Robust Device-Free Human Identification via WiFi and Vision Multimodal Learning

Lang Deng, Jianfei Yang, Shenghai Yuan, Han Zou, Chris Xiaoxuan Lu, Lihua Xie

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 12 pages, 8 figures, accepted by IEEE Internet of Things Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.03666 2022-08-19 cs.MM cs.CV cs.HC 81%

See What You See: Self-supervised Cross-modal Retrieval of Visual Stimuli from Brain Activity

Zesheng Ye, Lina Yao, Yu Zhang, Sylvia Gustin

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.03969 2022-07-05 cs.LG cs.CL cs.CV 81%

Multimodal Representations Learning Based on Mutual Information Maximization and Minimization and Identity Embedding for Multimodal Sentiment Analysis

Jiahao Zheng, Sen Zhang, Xiaoping Wang, Zhigang Zeng

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07190 2022-06-16 cs.CL cs.AI cs.LG 81%

Codec at SemEval-2022 Task 5: Multi-Modal Multi-Transformer Misogynous Meme Classification Framework

Ahmed Mahran, Carlo Alessandro Borella, Konstantinos Perifanos

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted for publication at the 16th International Workshop on Semantic Evaluation, Task 5: MAMI - Multimedia Automatic Misogyny Identification co-located with NAACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.14643 2022-05-31 cs.CV cs.AI cs.LG 81%

Micro-Expression Recognition Based on Attribute Information Embedding and Cross-modal Contrastive Learning

Yanxin Song, Jianzong Wang, Tianbo Wu, Zhangcheng Huang, Jing Xiao

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments This paper has been accepted by IJCNN2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08365 2022-05-18 cs.LG cs.AI cs.CV eess.IV 81%

Deep Supervised Information Bottleneck Hashing for Cross-modal Retrieval based Computer-aided Diagnosis

Yufeng Shi, Shuhuang Chen, Xinge You, Qinmu Peng, Weihua Ou, Yue Zhao

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 7 pages, 1 figure

Journal ref The AAAI-22 Workshop on Information Theory for Deep Learning (IT4DL).2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09860 2022-04-22 cs.CV cs.IR cs.MM 81%

Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information

Zhiqiang Yuan, Wenkai Zhang, Changyuan Tian, Xuee Rong, Zhengyuan Zhang, Hongqi Wang, Kun Fu, Xian Sun

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Journal ref in IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-16, 2022, Art no. 5620616

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.07668 2022-04-20 cs.CV cs.CL 81%

Dual-Key Multimodal Backdoors for Visual Question Answering

Matthew Walmer, Karan Sikka, Indranil Sur, Abhinav Shrivastava, Susmit Jha

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Published as conference paper at CVPR 2022. 22 pages, 11 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.07302 2022-04-18 cs.CV cs.CL 81%

Improving Cross-Modal Understanding in Visual Dialog via Contrastive Learning

Feilong Chen, Xiuyi Chen, Shuang Xu, Bo Xu

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.10852 2022-03-22 cs.CV cs.AI 81%

Multi-modal learning for predicting the genotype of glioma

Yiran Wei, Xi Chen, Lei Zhu, Lipei Zhang, Carola-Bibiane Schönlieb, Stephen J. Price, Chao Li

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏