arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3475 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3475 篇

2301.04742 2023-01-13 cs.CV 79%

HADA: A Graph-based Amalgamation Framework in Image-text Retrieval

Manh-Duy Nguyen, Binh T. Nguyen, Cathal Gurrin

专题命中 跨模态检索 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.04948 2022-12-19 cs.CV 79%

Transformer-based Cross-Modal Recipe Embeddings with Large Batch Training

Jing Yang, Junwen Chen, Keiji Yanai

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at MMM2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16208 2022-12-09 cs.CV 79%

SLAN: Self-Locator Aided Network for Cross-Modal Understanding

Jiang-Tian Zhai, Qi Zhang, Tong Wu, Xing-Yu Chen, Jiang-Jiang Liu, Bo Ren, Ming-Ming Cheng

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.01612 2022-12-06 cs.CL 79%

Named Entity and Relation Extraction with Multi-Modal Retrieval

Xinyu Wang, Jiong Cai, Yong Jiang, Pengjun Xie, Kewei Tu, Wei Lu

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

Comments Findings of EMNLP 2022. Code is publicly available at http://github.com/modelscope/adaseq/examples/MoRe

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.16262 2022-11-30 eess.IV cs.CV 79%

Is Image-to-Image Translation the Panacea for Multimodal Image Registration? A Comparative Study

Jiahao Lu, Johan Öfverstedt, Joakim Lindblad, Nataša Sladoje

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 37 pages, 10 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.12032 2022-11-24 cs.CV 79%

PointCMC: Cross-Modal Multi-Scale Correspondences Learning for Point Cloud Understanding

Honggu Zhou, Xiaogang Peng, Jiawei Mao, Zizhao Wu, Ming Zeng

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments In order to revise the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.11256 2022-11-22 cs.CL 79%

UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

Guimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu, Yuchuan Wu, Yongbin Li

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to EMNLP 2022 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.01884 2022-11-08 cs.LG cs.AI 79%

Graph Neural Networks for Multimodal Single-Cell Data Integration

Hongzhi Wen, Jiayuan Ding, Wei Jin, Yiqi Wang, Yuying Xie, Jiliang Tang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by KDD 2022 Applied Data Science Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10486 2022-10-20 cs.CV cs.LG 79%

Cross-Modal Fusion Distillation for Fine-Grained Sketch-Based Image Retrieval

Abhra Chaudhuri, Massimiliano Mancini, Yanbei Chen, Zeynep Akata, Anjan Dutta

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments British Machine Vision Conference (BMVC) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08908 2022-10-18 cs.CV 79%

Cross-modal Semantic Enhanced Interaction for Image-Sentence Retrieval

Xuri Ge, Fuhai Chen, Songpei Xu, Fuxiang Tao, Joemon M. Jose

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments accepted to WACV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04341 2022-10-11 cs.CV 79%

ConTra: (Con)text (Tra)nsformer for Cross-Modal Video Retrieval

Adriano Fragomeni, Michael Wray, Dima Damen

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted in ACCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.01416 2022-09-07 cs.AI cs.DB 79%

MMKGR: Multi-hop Multi-modal Knowledge Graph Reasoning

Shangfei Zheng, Weiqing Wang, Jianfeng Qu, Hongzhi Yin, Wei Chen, Lei Zhao

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02532 2022-08-05 cs.CL 79%

Prompt Tuning for Generative Multimodal Pretrained Models

Hao Yang, Junyang Lin, An Yang, Peng Wang, Chang Zhou, Hongxia Yang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.14656 2022-08-01 cs.CV cs.LG 79%

Multimodal SuperCon: Classifier for Drivers of Deforestation in Indonesia

Bella Septina Ika Hartanti, Valentino Vito, Aniati Murni Arymurthy, Andie Setiyoko

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.14428 2022-08-01 cs.CV 79%

Paired Cross-Modal Data Augmentation for Fine-Grained Image-to-Text Retrieval

Hao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan Miao

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12650 2022-07-27 cs.CV cs.IR 79%

Asymmetric Scalable Cross-modal Hashing

Wenyun Li, Chi-Man Pun

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01987 2022-07-06 cs.CV 79%

Open-Vocabulary 3D Detection via Image-level Class and Debiased Cross-modal Contrastive Learning

Yuheng Lu, Chenfeng Xu, Xiaobao Wei, Xiaodong Xie, Masayoshi Tomizuka, Kurt Keutzer, Shanghang Zhang

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.02770 2022-06-07 cs.CV 79%

Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts

Basil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton, Neil Houlsby

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.02343 2022-06-07 cs.CV 79%

Contrastive Graph Multimodal Model for Text Classification in Videos

Ye Liu, Changchong Lu, Chen Lin, Di Yin, Bo Ren

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.13656 2022-04-29 cs.CV 79%

Unsupervised Multi-Modal Medical Image Registration via Discriminator-Free Image-to-Image Translation

Zekang Chen, Jia Wei, Rui Li

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted in IJCAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10931 2022-04-26 cs.CL 79%

MCSE: Multimodal Contrastive Learning of Sentence Embeddings

Miaoran Zhang, Marius Mosbach, David Ifeoluwa Adelani, Michael A. Hedderich, Dietrich Klakow

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by NAACL 2022 main conference (short paper), 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08707 2022-04-20 cs.CV 79%

Unsupervised Contrastive Hashing for Cross-Modal Retrieval in Remote Sensing

Georgii Mikriukov, Mahdyar Ravanbakhsh, Begüm Demir

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments Our code is publicly available at https://git.tu-berlin.de/rsim/duch. arXiv admin note: substantial text overlap with arXiv:2201.08125

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05845 2022-04-13 cs.CV cs.LG 79%

Probabilistic Compositional Embeddings for Multimodal Image Retrieval

Andrei Neculai, Yanbei Chen, Zeynep Akata

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments CVPR2022 MULA workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03290 2022-04-04 cs.CV 79%

MC-LCR: Multi-modal contrastive classification by locally correlated representations for effective face forgery detection

Gaojian Wang, Qian Jiang, Xin Jin, Wei Li, Xiaohui Cui

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments 20 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.00680 2022-03-25 cs.CV 79%

CrossPoint: Self-Supervised Cross-Modal Contrastive Learning for 3D Point Cloud Understanding

Mohamed Afham, Isuru Dissanayake, Dinithi Dissanayake, Amaya Dharmasiri, Kanchana Thilakarathna, Ranga Rodrigo

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04260 2022-03-25 cs.MM cs.LG 79%

Online Enhanced Semantic Hashing: Towards Effective and Efficient Retrieval for Streaming Multi-Modal Data

Xiao-Ming Wu, Xin Luo, Yu-Wei Zhan, Chen-Lu Ding, Zhen-Duo Chen, Xin-Shun Xu

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.MM

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.14103 2022-03-22 cs.CV cs.LG 79%

Discriminative Semantic Transitive Consistency for Cross-Modal Learning

Kranti Kumar Parida, Gaurav Sharma

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.06920 2022-03-15 eess.IV cs.CV 79%

DS3-Net: Difficulty-perceived Common-to-T1ce Semi-Supervised Multimodal MRI Synthesis Network

Ziqi Huang, Li Lin, Pujin Cheng, Kai Pan, Xiaoying Tang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.01153 2022-03-03 cs.RO cs.AI 79%

InsertionNet 2.0: Minimal Contact Multi-Step Insertion Using Multimodal Multiview Sensory Input

Oren Spector, Vladimir Tchuiev, Dotan Di Castro

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to ICRA 2022, InsertionNet 1.0 : https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=9420246

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.10232 2022-02-22 cs.IR cs.CV cs.LG 79%

Efficient Cross-Modal Retrieval via Deep Binary Hashing and Quantization

Yang Shi, Young-joo Chung

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at BMVC 2021

Journal ref BMVC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏