arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3484 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3484 篇

2412.10761 2024-12-17 cs.CV cs.AI 62%

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation

Yang Yang, Wenjuan Xi, Luping Zhou, Jinhui Tang

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13282 2024-12-17 cs.CV cs.MM 62%

Wills Aligner: Multi-Subject Collaborative Brain Visual Decoding

Guangyin Bao, Qi Zhang, Zixuan Gong, Jialei Zhou, Wei Fan, Kun Yi, Usman Naseem, Liang Hu, Duoqian Miao

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments AAAI 2025, 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10348 2024-12-16 cs.CV cs.AI 62%

A dual contrastive framework

Yuan Sun, Zhao Zhang, Jorge Ortiz

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10155 2024-12-16 cs.CV cs.AI 62%

WordVIS: A Color Worth A Thousand Words

Umar Khan, Saifullah, Stefan Agne, Andreas Dengel, Sheraz Ahmed

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08231 2024-12-12 cs.CV cs.AI 62%

Dynamic Modality-Camera Invariant Clustering for Unsupervised Visible-Infrared Person Re-identification

Yiming Yang, Weipeng Hu, Haifeng Hu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12610 2024-12-10 cs.MM cs.SD eess.AS 62%

Emotion-Aligned Contrastive Learning Between Images and Music

Shanti Stewart, Kleanthis Avramidis, Tiantian Feng, Shrikanth Narayanan

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.MM、eess.AS

Comments Published at ICASSP 2024. Code: https://github.com/shantistewart/Emo-CLIM

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03957 2024-12-06 cs.CV cs.AI 62%

A Framework For Image Synthesis Using Supervised Contrastive Learning

Yibin Liu, Jianyu Zhang, Li Zhang, Shijian Li, Gang Pan

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00120 2024-12-03 cs.CV cs.AI 62%

Relation-Aware Meta-Learning for Zero-shot Sketch-Based Image Retrieval

Yang Liu, Jiale Du, Xinbo Gao, Jungong Han

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11875 2024-11-20 cs.IR cs.AI cs.CL q-bio.BM 62%

Exploring Optimal Transport-Based Multi-Grained Alignments for Text-Molecule Retrieval

Zijun Min, Bingshuai Liu, Liang Zhang, Jia Song, Jinsong Su, Song He, Xiaochen Bo

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL、cs.AI

Comments BIBM 2024 Regular Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09723 2024-11-18 cs.CV cs.AI 62%

Towards Neural Foundation Models for Vision: Aligning EEG, MEG, and fMRI Representations for Decoding, Encoding, and Modality Conversion

Matteo Ferrante, Tommaso Boccato, Grigorii Rashkov, Nicola Toschi

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17601 2024-11-18 cs.CV cs.AI 62%

CleanerCLIP: Fine-grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning

Yuan Xun, Siyuan Liang, Xiaojun Jia, Xinwei Liu, Xiaochun Cao

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06001 2024-11-13 cs.CV cs.MM 62%

Pseudo-triplet Guided Few-shot Composed Image Retrieval

Bohan Hou, Haoqiang Lin, Haokun Wen, Meng Liu, Mingzhu Xu, Xuemeng Song

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.MM

Comments 10pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11327 2024-11-05 cs.CL cs.AI 62%

Plug, Play, and Fuse: Zero-Shot Joint Decoding via Word-Level Re-ranking Across Diverse Vocabularies

Sai Koneru, Matthias Huck, Miriam Exel, Jan Niehues

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments WMT 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09952 2024-11-05 cs.CV cs.CL cs.LG 62%

BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval

Imanol Miranda, Ander Salaberria, Eneko Agirre, Gorka Azkune

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted to NeurIPS 24 Datasets and Benchmarks Track; Project page at: https://imirandam.github.io/BiVLC_project_page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14702 2024-11-01 cs.CV cs.AI 62%

G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality Models

Pengyue Jia, Yiding Liu, Xiaopeng Li, Yuhao Wang, Yantong Du, Xiao Han, Xuetao Wei, Shuaiqiang Wang, Dawei Yin, Xiangyu Zhao

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to NeurIPS2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08110 2024-10-31 cs.CL cs.CV 62%

Improving Hateful Meme Detection through Retrieval-Guided Contrastive Learning

Jingbiao Mei, Jinghong Chen, Weizhe Lin, Bill Byrne, Marcus Tomalin

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments ACL 2024 Main. The code is available from: https://github.com/JingbiaoMei/RGCL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21318 2024-10-30 cs.CV cs.AI 62%

Multi-path Exploration and Feedback Adjustment for Text-to-Image Person Retrieval

Bin Kang, Bin Chen, Junjie Wang, Yong Xu

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10184 2024-10-15 cs.CV cs.AI 62%

Eliminating the Language Bias for Visual Question Answering with fine-grained Causal Intervention

Ying Liu, Ge Bai, Chenji Lu, Shilong Li, Zhang Zhang, Ruifang Liu, Wenbin Guo

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

Journal ref 2024 IEEE International Conference on Multimedia and Expo (ICME), Niagara Falls, ON, Canada, 2024, pp. 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08593 2024-10-14 cs.CV cs.AI 62%

VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding

Houlun Chen, Xin Wang, Hong Chen, Zeyang Zhang, Wei Feng, Bin Huang, Jia Jia, Wenwu Zhu

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by 38th NeurIPS Datasets & Benchmarks Track (NeurIPS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17244 2024-10-11 cs.CL cond-mat.mtrl-sci cs.AI 62%

LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation

Yuan Chiang, Elvis Hsieh, Chia-Hong Chou, Janosh Riebesell

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 32 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06314 2024-10-10 cs.CV cs.CL 62%

Temporal Image Caption Retrieval Competition -- Description and Results

Jakub Pokrywka, Piotr Wierzchoń, Kornel Weryszko, Krzysztof Jassem

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Journal ref Proceedings of the 18th Conference on Computer Science and Intelligence Systems, M. Ganzha, L. Maciaszek, M. Paprzycki, D. Ślęzak (eds). ACSIS, Vol. 35, pages 1331-1336 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13119 2024-09-12 cs.CL cs.SD eess.AS 62%

Coarse-to-fine Alignment Makes Better Speech-image Retrieval

Lifeng Zhou, Yuke Li

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12023 2024-08-23 cs.LG cs.CL cs.CV 62%

Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them

Harish Haresamudram, Apoorva Beedu, Mashfiqui Rabbi, Sankalita Saha, Irfan Essa, Thomas Ploetz

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06259 2024-08-13 cs.CL cs.CV 62%

Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning

Yingjin Song, Denis Paperno, Albert Gatt

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 18 pages, 12 figures, accepted by INLG 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11317 2024-08-08 cs.CV cs.AI 62%

Improving Composed Image Retrieval via Contrastive Learning with Scaling Positives and Negatives

Zhangchi Feng, Richong Zhang, Zhijie Nie

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted to ACM MM 2024 Regular Papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13851 2024-07-22 cs.CV cs.LG cs.MM 62%

X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs

Sirnam Swetha, Jinyu Yang, Tal Neiman, Mamshad Nayeem Rizve, Son Tran, Benjamin Yao, Trishul Chilimbi, Mubarak Shah

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.MM

Comments Accepted at ECCV2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.07143 2024-07-16 cs.LG cs.AI cs.CV 62%

Reproducible scaling laws for contrastive language-image learning

Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, Jenia Jitsev

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.AI

Comments CVPR 2023. Version with minor extension. Original: CVPR2023/html/Cherti_Reproducible_Scaling_Laws_for_Contrastive_Language-Image_Learning_CVPR_2023_paper" target="_blank" rel="noopener">https://openaccess.thecvf.com/content/CVPR2023/html/Cherti_Reproducible_Scaling_Laws_for_Contrastive_Language-Image_Learning_CVPR_2023_paper

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 2818-2829

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06611 2024-07-10 cs.CV cs.AI 62%

CEIA: CLIP-Based Event-Image Alignment for Open-World Event-Based Understanding

Wenhao Xu, Wenming Weng, Yueyi Zhang, Zhiwei Xiong

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01408 2024-07-02 cs.CV cs.AI cs.LG 62%

Semantic Compositions Enhance Vision-Language Contrastive Learning

Maxwell Aladago, Lorenzo Torresani, Soroush Vosoughi

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18587 2024-06-28 cs.CV cs.AI 62%

Nomic Embed Vision: Expanding the Latent Space

Zach Nussbaum, Brandon Duderstadt, Andriy Mulyar

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏