arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3475 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3475 篇

2507.22938 2025-08-01 cs.CL cs.AI 76%

A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents

Sumit Soman, H. G. Ranjani, Sujoy Roychowdhury, Venkata Dharma Surya Narayana Sastry, Akshat Jain, Pranav Gangrade, Ayaaz Khan

机构 * Ericsson R&D Bangalore Karnataka India(爱立信研发部班加罗尔卡纳塔克邦印度)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CL、cs.AI

Comments Accepted for publication at the KDD 2025 Workshop on Structured Knowledge for Large Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15875 2025-07-23 cs.AI cs.MM 76%

Differential Multimodal Transformers

Jerry Li, Timothy Oh, Joseph Hoang, Vardhit Veeramachaneni

机构 * University of California, Riverside(加州大学河滨分校)

专题命中 跨模态检索 :multimodal(title);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14035 2025-06-18 cs.CV cs.AI 76%

SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement

Chelsi Jain, Yiran Wu, Yifan Zeng, Jiale Liu, S hengyu Dai, Zhenwen Shao, Qingyun Wu, Huazheng Wang

机构 * Oregon State University(俄勒冈州立大学) Pennsylvania State University(宾夕法尼亚州立大学) AG2AI, Inc.(AG2AI公司) Johnson & Johnson(强生公司)

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08854 2025-06-11 cs.CV cs.AI 76%

Spatial Transcriptomics Expression Prediction from Histopathology Based on Cross-Modal Mask Reconstruction and Contrastive Learning

Junzhuo Liu, Markus Eckstein, Zhixiang Wang, Friedrich Feuerhake, Dorit Merhof

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV、cs.AI

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08023 2025-06-11 q-bio.BM cs.AI cs.CE cs.CV cs.LG 76%

Aligning Proteins and Language: A Foundation Model for Protein Retrieval

Qifeng Wu, Zhengzhe Liu, Han Zhu, Yizhou Zhao, Daisuke Kihara, Min Xu

机构 * Carnegie Mellon University(卡内基梅隆大学) Purdue University(普渡大学)

专题命中 跨模态检索 :multimodal(abstract,comments);multimodal foundation model(abstract,comments);分类 cs.CV、cs.AI

Comments 4 pages for body, 3 pages for appendix, 11 figures. Accepted to CVPR 2025 Workshop on Multimodal Foundation Models for Biomedicine: Challenges and Opportunities(MMFM-BIOMED)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09519 2024-10-15 cs.CV cs.AI 76%

Pic@Point: Cross-Modal Learning by Local and Global Point-Picture Correspondence

Vencia Herzog, Stefan Suwelack

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV、cs.AI

Comments Accepted at ACML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08544 2024-08-19 cs.CV cs.MM 76%

Scaling up Multimodal Pre-training for Sign Language Understanding

Wengang Zhou, Weichao Zhao, Hezhen Hu, Zecheng Li, Houqiang Li

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.MM

Comments Sign language recognition; Sign language translation; Sign language retrieval

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.06357 2024-08-14 cs.CV cs.AI 76%

Algorithm Research of ELMo Word Embedding and Deep Learning Multimodal Transformer in Image Description

Xiaohan Cheng, Taiyuan Mei, Yun Zi, Qi Wang, Zijun Gao, Haowei Yang

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00599 2024-08-02 cs.CV cs.MM eess.IV 76%

Learned Compression of Point Cloud Geometry and Attributes in a Single Model through Multimodal Rate-Control

Michael Rudolph, Aron Riemenschneider, Amr Rizk

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.MM

Comments 20 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06201 2024-06-11 cs.CV cs.AI 76%

2DP-2MRC: 2-Dimensional Pointer-based Machine Reading Comprehension Method for Multimodal Moment Retrieval

Jiajun He, Tomoki Toda

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

Comments Accepted by INTERSPEECH 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08851 2024-03-15 astro-ph.IM cs.CL cs.CV cs.IR cs.LG 76%

PAPERCLIP: Associating Astronomical Observations and Natural Language with Multi-Modal Models

Siddharth Mishra-Sharma, Yiding Song, Jesse Thaler

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV、cs.CL

Comments 17+6 pages, 3+1 figures, 5+2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12846 2024-02-21 cs.CV cs.AI 76%

ConVQG: Contrastive Visual Question Generation with Multimodal Guidance

Li Mi, Syrielle Montariol, Javiera Castillo-Navarro, Xianjie Dai, Antoine Bosselut, Devis Tuia

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

Comments AAAI 2024. Project page at https://limirs.github.io/ConVQG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.12258 2023-08-29 cs.SD cs.CL cs.IR cs.LG eess.AS 76%

Data leakage in cross-modal retrieval training: A case study

Benno Weck, Xavier Serra

专题命中 跨模态检索 :cross-modal(title);分类 cs.CL、eess.AS

Comments 5 pages. Accepted at ICASSP2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.12996 2023-07-26 cs.LG cs.AI cs.CL cs.IR q-bio.QM 76%

Extracting Molecular Properties from Natural Language with Multimodal Contrastive Learning

Romain Lacombe, Andrew Gaut, Jeff He, David Lüdeke, Kateryna Pistunova

专题命中 跨模态检索 :multimodal(title);分类 cs.CL、cs.AI

Comments 2023 ICML Workshop on Computational Biology

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10893 2023-04-24 cs.CV cs.MM 76%

FindVehicle and VehicleFinder: A NER dataset for natural language-based vehicle retrieval and a keyword-based cross-modal vehicle retrieval system

Runwei Guan, Ka Lok Man, Feifan Chen, Shanliang Yao, Rongsheng Hu, Xiaohui Zhu, Jeremy Smith, Eng Gee Lim, Yutao Yue

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12413 2022-12-27 cs.CV cs.AI 76%

Segmentation of Parotid Gland Tumors Using Multimodal MRI and Contrastive Learning

Zi'an Xu, Yin Dai, Fayu Liu, Boyuan Wu, Weibing Chen, Lifu Shi

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14395 2022-10-27 cs.CV cs.CL cs.LG 76%

IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text

Seungwhan Moon, Andrea Madotto, Zhaojiang Lin, Alireza Dirafzoon, Aparajita Saraf, Amy Bearman, Babak Damavandi

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.01445 2022-03-07 cs.CV cs.CL cs.LG eess.IV 76%

LILE: Look In-Depth before Looking Elsewhere -- A Dual Attention Network using Transformers for Cross-Modal Information Retrieval in Histopathology Archives

Danial Maleki, H. R Tizhoosh

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.12932 2019-10-01 cs.CV cs.HC cs.IR cs.MM 76%

BUDA.ART: A Multimodal Content-Based Analysis and Retrieval System for Buddha Statues

Benjamin Renoust, Matheus Oliveira Franca, Jacob Chan, Van Le, Ayaka Uesaka, Yuta Nakashima, Hajime Nagahara, Jueren Wang, Yutaka Fujioka

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.MM

Comments Demo video at: https://www.youtube.com/watch?v=3XJvLjSWieY

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.11299 2019-05-15 cs.CV cs.CL 76%

Image search using multilingual texts: a cross-modal learning approach between image and text

Maxime Portaz, Hicham Randrianarivo, Adrien Nivaggioli, Estelle Maudet, Christophe Servan, Sylvain Peyronnet

专题命中 跨模态检索 :cross-modal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.00860 2017-07-27 cs.LG cs.AI cs.CV 76%

Conditional generation of multi-modal data using constrained embedding space mapping

Subhajit Chaudhury, Sakyasingha Dasgupta, Asim Munawar, Md. A. Salam Khan, Ryuki Tachibana

专题命中 跨模态检索 :multi-modal(title);分类 cs.CV、cs.AI

Comments 7 pages, 4 figures, ICML 2017 Workshop on Implicit Models

详情

展开后加载摘要…

URL PDF HTML 收藏
1509.07831 2017-05-18 cs.RO cs.AI cs.CV cs.LG 76%

Deep Multimodal Embedding: Manipulating Novel Objects with Point-clouds, Language and Trajectories

Jaeyong Sung, Ian Lenz, Ashutosh Saxena

专题命中 跨模态检索 :multimodal(title);分类 cs.CV、cs.AI

Comments IEEE International Conference on Robotics and Automation (ICRA), 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.06078 2016-04-15 cs.CV cs.CL cs.LG 76%

Learning Deep Structure-Preserving Image-Text Embeddings

Liwei Wang, Yin Li, Svetlana Lazebnik

专题命中 跨模态检索 :image-text(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04812 2026-03-03 cs.CV cs.AI cs.CL cs.LG 75%

LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning

LLaVE: 大规模语言和视觉嵌入模型与基于难度加权的对比学习

Zhibin Lan, Liqiang Niu, Fandong Meng, Jie Zhou, Jinsong Su

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) Pattern Recognition Center, WeChat AI, Tencent Inc, China(腾讯人工智能研究院) Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 LLaVE通过基于难度加权的对比学习提升多模态嵌入模型性能,实现SOTA表现和强泛化能力。

Comments Accepted by Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17960 2025-10-22 astro-ph.IM astro-ph.CO 75%

AION-1: Omnimodal Foundation Model for Astronomical Sciences

Liam Parker, Francois Lanusse, Jeff Shen, Ollie Liu, Tom Hehir, Leopoldo Sarra, Lucas Meyer, Micah Bowles, Sebastian Wagner-Carena, Helen Qu, Siavash Golkar, Alberto Bietti, Hatim Bourfoune, Nathan Casserau, Pierre Cornette, Keiya Hirashima, Geraud Krawezik, Ruben Ohana, Nicholas Lourie, Michael McCabe, Rudy Morel, Payel Mukhopadhyay, Mariel Pettee, Bruno Regaldo-Saint Blancard, Kyunghyun Cho, Miles Cranmer, Shirley Ho

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract)

Comments Accepted at Neural Information Processing Systems (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23379 2025-10-20 cs.CL cs.AI cs.CV 75%

CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding

Xi Zhang, Zaiqiao Meng, Jake Lever, Edmond S. L. Ho

机构 * School of Computing Science, University of Glasgow(计算科学学院,格拉斯哥大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Preprint, 27 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12131 2025-03-18 cs.CV cs.AI cs.LG cs.SD eess.AS 75%

DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap

Shentong Mo, Zehua Chen, Fan Bao, Jun Zhu

专题命中 跨模态检索 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15052 2025-01-28 cs.CV cs.AI cs.MM 75%

Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval

Bingjun Luo, Jinpeng Wang, Wang Zewen, Junjie Zhu, Xibin Zhao

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00986 2022-12-06 cs.CV cs.AI cs.CL 75%

Masked Contrastive Pre-Training for Efficient Video-Text Retrieval

Fangxun Shu, Biaolong Chen, Yue Liao, Shuwen Xiao, Wenyu Sun, Xiaobo Li, Yousong Zhu, Jinqiao Wang, Si Liu

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.07194 2020-08-05 cs.CL cs.CV cs.IR cs.LG cs.MM 75%

Recommending Themes for Ad Creative Design via Visual-Linguistic Representations

Yichao Zhou, Shaunak Mishra, Manisha Verma, Narayan Bhamidipati, Wei Wang

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments 7 pages, 8 figures, 2 tables, accepted by The Web Conference 2020

详情

展开后加载摘要…

URL PDF HTML 收藏