arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3475 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3475 篇

2203.00187 2024-02-15 cs.RO cs.CV 79%

Robots Autonomously Detecting People: A Multimodal Deep Contrastive Learning Method Robust to Intraclass Variations

Angus Fung, Beno Benhabib, Goldie Nejat

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07736 2024-02-13 cs.IR cs.MM 79%

Multimodal Learned Sparse Retrieval for Image Suggestion

Thong Nguyen, Mariya Hendriksen, Andrew Yates

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.MM

Comments 5 pages, TREC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07016 2024-02-13 cs.AI 79%

REALM: RAG-Driven Enhancement of Multimodal Electronic Health Records Analysis via Large Language Models

Yinghao Zhu, Changyu Ren, Shiyun Xie, Shukai Liu, Hangyuan Ji, Zixiang Wang, Tao Sun, Long He, Zhoujun Li, Xi Zhu, Chengwei Pan

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17954 2024-02-09 cs.CV 79%

Transformer-empowered Multi-modal Item Embedding for Enhanced Image Search in E-Commerce

Chang Liu, Peng Hou, Anxiang Zeng, Han Yu

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by IAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.10789 2024-01-31 cs.LG cs.CL q-bio.QM stat.ML 79%

Multi-modal Molecule Structure-text Model for Text-based Retrieval and Editing

Shengchao Liu, Weili Nie, Chengpeng Wang, Jiarui Lu, Zhuoran Qiao, Ling Liu, Jian Tang, Chaowei Xiao, Anima Anandkumar

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16304 2024-01-30 cs.CV cs.IR cs.LG 79%

Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder

Zheyuan Liu, Weixuan Sun, Damien Teney, Stephen Gould

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at TMLR, 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12225 2024-01-24 cs.CV cs.LG 79%

Multimodal Data Curation via Object Detection and Filter Ensembles

Tzu-Heng Huang, Changho Shin, Sui Jiet Tay, Dyah Adila, Frederic Sala

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Appeared in the Workshop of Towards the Next Generation of Computer Vision Datasets (TNGCV) on ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11847 2024-01-23 cs.CV 79%

SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning

Hao Chen, Jiaze Wang, Ziyu Guo, Jinpeng Li, Donghao Zhou, Bian Wu, Chenyong Guan, Guangyong Chen, Pheng-Ann Heng

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08619 2024-01-18 cs.LG cs.AI 79%

MATE-Pred: Multimodal Attention-based TCR-Epitope interaction Predictor

Etienne Goffinet, Raghvendra Mall, Ankita Singh, Rahul Kaushik, Filippo Castiglione

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.AI

Comments Patent pending: U.S. Provisional Application No. 63/603,952

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.17093 2024-01-17 cs.CV 79%

Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval

Hao Li, Jingkuan Song, Lianli Gao, Xiaosu Zhu, Heng Tao Shen

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted to NeurIPS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02582 2024-01-08 cs.CV 79%

CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Daoan Zhang, Junming Yang, Hanjia Lyu, Zijian Jin, Yuan Yao, Mingkai Chen, Jiebo Luo

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01495 2024-01-04 cs.CL 79%

A Two-Stage Multimodal Emotion Recognition Model Based on Graph Contrastive Learning

Wei Ai, FuChen Zhang, Tao Meng, YunTao Shou, HongEn Shao, Keqin Li

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16741 2024-01-03 cs.LG cs.AI cs.HC 79%

Multi-Modal Financial Time-Series Retrieval Through Latent Space Projections

Tom Bamford, Andrea Coletta, Elizabeth Fons, Sriram Gopalakrishnan, Svitlana Vyetrenko, Tucker Balch, Manuela Veloso

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted to ICAIF 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15021 2023-12-27 cs.CL 79%

Towards a Unified Multimodal Reasoning Framework

Abhinav Arun, Dipendra Singh Mal, Mehul Soni, Tomohiro Sawada

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments 6 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10547 2023-12-15 cs.CV cs.CY 79%

Rethinking Multimodal Content Moderation from an Asymmetric Angle with Mixed-modality

Jialin Yuan, Ye Yu, Gaurav Mittal, Matthew Hall, Sandra Sajeev, Mei Chen

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00857 2023-12-05 cs.LG cs.AI cs.HC eess.SP 79%

Latent Space Explorer: Visual Analytics for Multimodal Latent Space Exploration

Bum Chul Kwon, Samuel Friedman, Kai Xu, Steven A Lubitz, Anthony Philippakis, Puneet Batra, Patrick T Ellinor, Kenney Ng

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.11009 2023-11-21 cs.CL 79%

Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimodal Emotion Recognition

Dongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu Okumura

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.04678 2023-11-14 cs.CV q-bio.QM 79%

Weakly supervised cross-modal learning in high-content screening

Watkinson Gabriel, Cohen Ethan, Bourriez Nicolas, Bendidi Ihab, Bollot Guillaume, Genovesio Auguste

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.03964 2023-11-08 cs.CV 79%

Enhancing Multimodal Compositional Reasoning of Visual Language Models with Generative Negative Mining

Ugur Sahin, Hang Li, Qadeer Khan, Daniel Cremers, Volker Tresp

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to WACV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.13214 2023-11-08 cs.LG cs.AI 79%

FedMEKT: Distillation-based Embedding Knowledge Transfer for Multimodal Federated Learning

Huy Q. Le, Minh N. H. Nguyen, Chu Myaet Thwal, Yu Qiao, Chaoning Zhang, Choong Seon Hong

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18770 2023-10-31 cs.IR cs.MM 79%

Leveraging Multimodal Features and Item-level User Feedback for Bundle Construction

Yunshan Ma, Xiaohao Liu, Yinwei Wei, Zhulin Tao, Xiang Wang, Tat-Seng Chua

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

Journal ref WSDM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.13265 2023-10-23 cs.CL 79%

MoqaGPT : Zero-Shot Multi-modal Open-domain Question Answering with Large Language Model

Le Zhang, Yihong Wu, Fengran Mo, Jian-Yun Nie, Aishwarya Agrawal

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted into EMNLP2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.11029 2023-10-19 cs.SD cs.IR eess.AS 79%

CLaMP: Contrastive Language-Music Pre-training for Cross-Modal Symbolic Music Information Retrieval

Shangda Wu, Dingyao Yu, Xu Tan, Maosong Sun

专题命中 跨模态检索 :cross-modal(title,abstract);分类 eess.AS

Comments 11 pages, 5 figures, 5 tables, accepted by ISMIR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.12403 2023-10-16 cs.CV cs.LG cs.RO 79%

ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal Embeddings

Arjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman, Dhruv Batra

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments code: https://github.com/gunagg/zson

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08032 2023-10-13 cs.AI 79%

Incorporating Domain Knowledge Graph into Multimodal Movie Genre Classification with Self-Supervised Attention and Contrastive Learning

Jiaqi Li, Guilin Qi, Chuanyi Zhang, Yongrui Chen, Yiming Tan, Chenlong Xia, Ye Tian

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07668 2023-10-12 cs.LG cs.AI 79%

GRaMuFeN: Graph-based Multi-modal Fake News Detection in Social Media

Makan Kananian, Fatima Badiei, S. AmirAli Gh. Ghahramani

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.08671 2023-10-10 cs.CR cs.AI 79%

Deep Cross-Modal Steganography Using Neural Representations

Gyojin Han, Dong-Jae Lee, Jiwan Hur, Jaehyun Choi, Junmo Kim

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.AI

Comments ICIP 2023 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.11880 2023-10-06 cs.MM 79%

Adaptive Marginalized Semantic Hashing for Unpaired Cross-Modal Retrieval

Kaiyi Luo, Chao Zhang, Huaxiong Li, Xiuyi Jia, Chunlin Chen

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11553 2023-09-26 cs.CV 79%

MuMUR : Multilingual Multimodal Universal Retrieval

Avinash Madasu, Estelle Aflalo, Gabriela Ben Melech Stan, Shachar Rosenman, Shao-Yen Tseng, Gedas Bertasius, Vasudev Lal

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments This is an extension of the previous MKTVR paper (for which you can find a reference here : https://dl.acm.org/doi/abs/10.1007/978-3-031-28244-7_42 or in a previous version on arxiv). This version was published to the Information Retrieval Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05032 2023-09-12 cs.CV 79%

Unified Contrastive Fusion Transformer for Multimodal Human Action Recognition

Kyoung Ok Yang, Junho Koh, Jun Won Choi

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏