arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3484 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3484 篇

2405.19547 2024-12-23 cs.LG cs.CV 74%

CLIPLoss and Norm-Based Data Selection Methods for Multimodal Contrastive Learning

Yiping Wang, Yifang Chen, Wendan Yan, Alex Fang, Wenjing Zhou, Kevin Jamieson, Simon Shaolei Du

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments This paper supercedes our previous VAS paper (arXiv:2402.02055). It's accepted by NeurIPS2024 as spotlight paper. DataComp benchmark: https://www.datacomp.ai/dcclip/leaderboard.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07619 2024-12-11 cs.CL 74%

DRUM: Learning Demonstration Retriever for Large MUlti-modal Models

Ellen Yi-Ge, Jiechao Gao, Wei Han, Wei Zhu

专题命中 跨模态检索 :multi-modal(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23437 2024-11-01 cs.LG cs.CL cs.IR 74%

Mind the Gap: A Generalized Approach for Cross-Modal Embedding Alignment

Arihan Yadav, Alan McMillan

专题命中 跨模态检索 :cross-modal(title);分类 cs.CL

Comments 18 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06827 2024-09-12 cs.CV 74%

Cross-Modal Self-Supervised Learning with Effective Contrastive Units for LiDAR Point Clouds

Mu Cai, Chenxu Luo, Yong Jae Lee, Xiaodong Yang

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

Comments IROS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07925 2024-07-12 cs.IR cs.AI cs.SI 74%

Enhancing Social Media Personalization: Dynamic User Profile Embeddings and Multimodal Contextual Analysis Using Transformer Models

Pranav Vachharajani

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

Comments 21 pages, 13 figures. Mentor: Prof Pritam Ranjan

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17615 2024-06-26 cs.SE cs.AI cs.LG 74%

Aligning Programming Language and Natural Language: Exploring Design Choices in Multi-Modal Transformer-Based Embedding for Bug Localization

Partha Chakraborty, Venkatraman Arumugam, Meiyappan Nagappan

专题命中 跨模态检索 :multi-modal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12024 2024-01-23 cs.RO cs.AI cs.LG 74%

Multimodal Visual-Tactile Representation Learning through Self-Supervised Contrastive Pre-Training

Vedant Dave, Fotios Lygerakis, Elmar Rueckert

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03043 2024-01-09 cs.CV 74%

Learning Multimodal Volumetric Features for Large-Scale Neuron Tracing

Qihua Chen, Xuejin Chen, Chenxuan Wang, Yixiong Liu, Zhiwei Xiong, Feng Wu

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments 9 pages, 6 figures, AAAI 2024 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00693 2023-12-12 cs.AI 74%

On Task-personalized Multimodal Few-shot Learning for Visually-rich Document Entity Retrieval

Jiayi Chen, Hanjun Dai, Bo Dai, Aidong Zhang, Wei Wei

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

Comments Paper published at Findings of the Association for Computational Linguistics: EMNLP, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06286 2023-11-07 eess.SP cs.CV 74%

Automated Cardiovascular Record Retrieval by Multimodal Learning between Electrocardiogram and Clinical Report

Jielin Qiu, Jiacheng Zhu, Shiqi Liu, William Han, Jingqi Zhang, Chaojing Duan, Michael Rosenberg, Emerson Liu, Douglas Weber, Ding Zhao

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments Accepted to the ML4H 2023 Proceedings track

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.16386 2023-09-01 cs.CV 74%

RGB-T Tracking via Multi-Modal Mutual Prompt Learning

Yang Luo, Xiqing Guo, Hui Feng, Lei Ao

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

Comments 9 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14523 2023-07-28 cs.CV 74%

Towards multi-modal anatomical landmark detection for ultrasound-guided brain tumor resection with contrastive learning

Soorena Salari, Amirhossein Rasoulian, Hassan Rivaz, Yiming Xiao

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

Comments Accepted in MICCAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13805 2023-05-24 cs.CL 74%

Towards Zero-shot Relation Extraction in Web Mining: A Multimodal Approach with Relative XML Path

Zilong Wang, Jingbo Shang

专题命中 跨模态检索 :multimodal(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08660 2023-04-19 cs.RO cs.AI 74%

(LC)$^2$: LiDAR-Camera Loop Constraints For Cross-Modal Place Recognition

Alex Junho Lee, Seungwon Song, Hyungtae Lim, Woojoo Lee, Hyun Myung

专题命中 跨模态检索 :cross-modal(title);分类 cs.AI

Comments 8 pages, 11 figures, Accepted to IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.15770 2023-03-29 eess.IV cs.CV physics.med-ph 74%

DDMM-Synth: A Denoising Diffusion Model for Cross-modal Medical Image Synthesis with Sparse-view Measurement Embedding

Xiaoyue Li, Kai Shang, Gaoang Wang, Mark D. Butala

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

Comments llncs.cls v2.20,12 pages with 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.01824 2022-11-04 cs.CL 74%

Human in the loop approaches in multi-modal conversational task guidance system development

Ramesh Manuvinakurike, Sovan Biswas, Giuseppe Raffa, Richard Beckwith, Anthony Rhodes, Meng Shi, Gesem Gudino Mejia, Saurav Sahay, Lama Nachman

专题命中 跨模态检索 :multi-modal(title);分类 cs.CL

Comments SCAI @ SIGIR

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03838 2022-10-11 cs.CV 74%

Learning to embed semantic similarity for joint image-text retrieval

Noam Malali, Yosi Keller

专题命中 跨模态检索 :image-text(title);分类 cs.CV

Comments in IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.02329 2022-09-07 cs.CV cs.LG 74%

Multimodal contrastive learning for remote sensing tasks

Umangi Jain, Alex Wilson, Varun Gulshan

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07734 2022-07-26 q-bio.GN cs.AI cs.GL 74%

COEM: Cross-Modal Embedding for MetaCell Identification

Haiyi Mao, Minxue Jia, Jason Xiaotian Dou, Haotian Zhang, Panayiotis V. Benos

专题命中 跨模态检索 :cross-modal(title);分类 cs.AI

Comments 5 pages, 2 figures, ICML workshop on computational biology

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.08533 2022-06-14 cs.CV 74%

Eliminate Deviation with Deviation for Data Augmentation and a General Multi-modal Data Learning Method

Yunpeng Gong, Liqing Huang, Lifei Chen

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.04426 2022-05-26 cs.CV 74%

Deep Feature Rotation for Multimodal Image Style Transfer

Son Truong Nguyen, Nguyen Quang Tuyen, Nguyen Hong Phuc

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments Accepted to NICS'21

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.15914 2022-05-24 cs.CV 74%

Tasting the cake: evaluating self-supervised generalization on out-of-distribution multimodal MRI data

Alex Fedorov, Eloy Geenjaar, Lei Wu, Thomas P. DeRamus, Vince D. Calhoun, Sergey M. Plis

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments Presented as a RobustML workshop paper at ICLR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04021 2022-05-02 cs.CV 74%

Supervised Contrastive Learning for Detecting Anomalous Driving Behaviours from Multimodal Videos

Shehroz S. Khan, Ziting Shen, Haoying Sun, Ax Patel, Ali Abedi

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments 8 pages, 2 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.07052 2022-04-15 cs.CV 74%

CroCo: Cross-Modal Contrastive learning for localization of Earth Observation data

Wei-Hsin Tseng, Hoàng-Ân Lê, Alexandre Boulch, Sébastien Lefèvre, Dirk Tiede

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

Comments Accepted for publication in the ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences (online from July 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.05431 2021-11-11 cs.LG cs.AI 74%

Multi-Task Prediction of Clinical Outcomes in the Intensive Care Unit using Flexible Multimodal Transformers

Benjamin Shickel, Patrick J. Tighe, Azra Bihorac, Parisa Rashidi

专题命中 跨模态检索 :multimodal(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.11592 2021-10-25 cs.CV cs.IR 74%

Learning Text-Image Joint Embedding for Efficient Cross-Modal Retrieval with Deep Feature Engineering

Zhongwei Xie, Ling Liu, Yanzhao Wu, Luo Zhong, Lin Li

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

Comments accepted by ACM Transactions on Information Systems(TOIS). arXiv admin note: text overlap with arXiv:2108.00705, arXiv:2108.03788

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.01654 2021-08-12 cs.CV 74%

Ask&Confirm: Active Detail Enriching for Cross-Modal Retrieval with Partial Query

Guanyu Cai, Jun Zhang, Xinyang Jiang, Yifei Gong, Lianghua He, Fufu Yu, Pai Peng, Xiaowei Guo, Feiyue Huang, Xing Sun

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV

Comments Accepted by ICCV2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.13711 2021-06-28 cs.IR cs.CL 74%

Multimodal Emergent Fake News Detection via Meta Neural Process Networks

Yaqing Wang, Fenglong Ma, Haoyu Wang, Kishlay Jha, Jing Gao

专题命中 跨模态检索 :multimodal(title);分类 cs.CL

Comments accepted by KDD 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.08665 2021-05-19 cs.LG cs.CV 74%

A multimodal deep learning framework for scalable content based visual media retrieval

Ambareesh Ravi, Amith Nandakumar

专题命中 跨模态检索 :multimodal(title);分类 cs.CV

Comments Paper pertaining to a course project

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.12482 2021-02-23 cs.CV cs.SI 74%

On the Limits to Multi-Modal Popularity Prediction on Instagram -- A New Robust, Efficient and Explainable Baseline

Christoffer Riis, Damian Konrad Kowalczyk, Lars Kai Hansen

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV

Comments Presented at ICAART 2021

Journal ref Proceedings of the 13th International Conference on Agents and Artificial Intelligence - Volume 2: ICAART, ISBN 978-989-758-484-8, pages 1200-1209, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏