arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3475 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3475 篇

2405.07460 2025-08-28 cs.LG cs.AI cs.DB 79%

HoneyBee: A Scalable Modular Framework for Creating Multimodal Oncology Datasets with Foundational Embedding Models

Aakash Tripathi, Asim Waqas, Matthew B. Schabath, Yasin Yilmaz, Ghulam Rasool

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18132 2025-08-26 cs.IR cs.AI cs.LG 79%

Test-Time Scaling Strategies for Generative Retrieval in Multimodal Conversational Recommendations

Hung-Chun Hsu, Yuan-Ching Kuo, Chao-Han Huck Yang, Szu-Wei Fu, Hanrong Ye, Hongxu Yin, Yu-Chiang Frank Wang, Ming-Feng Tsai, Chuan-Ju Wang

机构 * Research Center for Information Technology Innovation, Academia Sinica(资讯科技创新研究所以) NVIDIA(NVIDIA公司) Department of Computer Science, National Chengchi University(国立政治大学计算机科学系)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17044 2025-08-26 cs.CV cs.RO 79%

M3DMap: Object-aware Multimodal 3D Mapping for Dynamic Environments

Dmitry Yudin

机构 * Moscow Institute of Physics and Technology(莫斯科物理技术学院) AIRI

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 29 pages, 3 figures, 13 tables. Preprint of the accepted article in Optical Memory and Neural Network Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01273 2025-08-26 cs.IR cs.MM 79%

SoccerRAG: Multimodal Soccer Information Retrieval via Natural Queries

Aleksander Theo Strand, Sushant Gautam, Cise Midoglu, Pål Halvorsen

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

Comments accepted to CBMI 2024 as a regular paper; https://github.com/simula/soccer-rag

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16882 2025-08-26 eess.IV cs.CV 79%

Multimodal Medical Endoscopic Image Analysis via Progressive Disentangle-aware Contrastive Learning

Junhao Wu, Yun Li, Junhao Li, Jingliang Bian, Xiaomao Fan, Wenbin Lei, Ruxin Wang

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) First Affiliated Hospital, Sun Yat-sen University(中山大学第一附属医院) College of Big Data and Internet, Shenzhen Technology University(深圳技术大学大数据与互联网学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments 12 pages,6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10264 2025-08-25 cs.LG cs.AI cs.IR 79%

Order-Preserving Dimension Reduction for Multimodal Semantic Embedding

Chengyu Gong, Gefei Shen, Luanzheng Guo, Nathan Tallent, Dongfang Zhao

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14515 2025-08-21 cs.IR cs.AI 79%

MISS: Multi-Modal Tree Indexing and Searching with Lifelong Sequential Behavior for Retrieval Recommendation

Chengcheng Guo, Junda She, Kuo Cai, Shiyao Wang, Qigen Hu, Qiang Luo, Kun Gai, Guorui Zhou

机构 * Kuaishou Inc.(快手公司)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

Comments CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22884 2025-08-20 cs.CV 79%

AutoComPose: Automatic Generation of Pose Transition Descriptions for Composed Pose Retrieval Using Multimodal LLMs

Yi-Ting Shen, Sungmin Eum, Doheon Lee, Rohit Shete, Chiao-Yi Wang, Heesung Kwon, Shuvra S. Bhattacharyya

机构 * University of Maryland, College Park(马里兰大学学院市分校) DEVCOM Army Research Laboratory(陆军研究实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12149 2025-08-19 cs.AI 79%

MOVER: Multimodal Optimal Transport with Volume-based Embedding Regularization

Haochen You, Baojing Liu

机构 * Columbia University(哥伦比亚大学) Hebei Institute of Communications(河北通信学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Accepted as a conference paper at CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17297 2025-08-08 cs.AI 79%

Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts

Zhenghao Liu, Xingsheng Zhu, Tianshuo Zhou, Xinyi Zhang, Xiaoyuan Yi, Yukun Yan, Ge Yu, Maosong Sun

机构 * Northeastern University, China(东北大学) Microsoft Research Asia(微软亚洲研究院) Tsinghua University(清华大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19264 2025-08-07 cs.CV 79%

SimMLM: A Simple Framework for Multi-modal Learning with Missing Modality

Sijie Li, Chen Chen, Jungong Han

机构 * School of Computer Science, University of Sheffield, UK(计算机科学学院,谢菲尔德大学)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV

Journal ref ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03494 2025-08-06 cs.CV 79%

Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval

Shreyank N Gowda, Xiaobo Jin, Christian Wagner

机构 * School of Computer Science, The University of Nottingham, NG8 1BB Nottingham, U.K.(计算机科学学院,诺丁汉大学) Department of Intelligent Science, Xi’an Jiaotong-Liverpool University, China, 215123.(智能科学系,西安交通大学利物浦大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00332 2025-08-04 cs.CL 79%

Improving Multimodal Contrastive Learning of Sentence Embeddings with Object-Phrase Alignment

Kaiyan Zhao, Zhongtao Miao, Yoshimasa Tsuruoka

机构 * The University of Tokyo(东京大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23188 2025-08-01 cs.CV 79%

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space

Shiyao Yu, Zi-An Wang, Kangning Yin, Zheng Tian, Mingyuan Zhang, Weixin Si, Shihao Zou

机构 * Southern University of Science and Technology(南方科技大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) ShanghaiTech University(上海科技大学) Nanyang Technological University(南洋理工大学) Faculty of Computer Science and Control Engineering, Shenzhen University of Advanced Technology(计算机科学与控制工程学院,深圳大学先进技术学院)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by IEEE TMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20259 2025-07-29 cs.CV 79%

L-MCAT: Unpaired Multimodal Transformer with Contrastive Attention for Label-Efficient Satellite Image Classification

Mitul Goswami, Mrinal Goswami

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20189 2025-07-29 eess.SP cs.AI cs.LG q-bio.NC 79%

NeuroCLIP: A Multimodal Contrastive Learning Method for rTMS-treated Methamphetamine Addiction Analysis

Chengkai Wang, Di Wu, Yunsheng Liao, Wenyao Zheng, Ziyi Zeng, Xurong Gao, Hemmings Wu, Zhoule Zhu, Jie Yang, Lihua Zhong, Weiwei Cheng, Yun-Hsuan Chen, Mohamad Sawan

机构 * CenBRAIN Neurotech Center of Excellence, School of Engineering, Westlake University(西溪大学工程学院先进神经技术中心) School of Data Science, Xiamen University Malaysia(马来西亚厦门大学数据科学学院) Department of Neurosurgery, Second Affiliated Hospital, School of Medicine, Zhejiang University(浙江大学医学院附属第二医院神经外科) Department of Education and Correction, Zhejiang Gongchen Compulsory Isolated Detoxification Center(浙江省公检强制隔离戒毒所教育矫正部) Zhejiang Liangzhu Compulsory Isolated Detoxification Center(浙江省良渚强制隔离戒毒所)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00970 2025-07-25 cs.MM 79%

Multimodal Fusion via Hypergraph Autoencoder and Contrastive Learning for Emotion Recognition in Conversation

Zijian Yi, Ziming Zhao, Zhishu Shen, Tiehua Zhang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

Comments Accepted by ACM MULTIMEDIA 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17533 2025-07-24 cs.CV 79%

Multi-modal Multi-task Pre-training for Improved Point Cloud Understanding

Liwen Liu, Weidong Yang, Lipeng Ma, Ben Fei

机构 * Fudan University(复旦大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07942 2025-07-22 cs.CV 79%

MARS: a Multimodal Alignment and Ranking System for Few-Shot Segmentation

Nico Catalano, Stefano Samele, Paolo Pertino, Matteo Matteucci

机构 * Politecnico di Milano(米兰理工大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14738 2025-07-22 cs.CV 79%

MultiRetNet: A Multimodal Vision Model and Deferral System for Staging Diabetic Retinopathy

Jeannie She, Katie Spivakovsky

机构 * MIT(麻省理工学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12998 2025-07-18 cs.CV cs.LG 79%

Differential-informed Sample Selection Accelerates Multimodal Contrastive Learning

Zihua Zhao, Feng Hong, Mengxi Chen, Pengyi Chen, Benyuan Liu, Jiangchao Yao, Ya Zhang, Yanfeng Wang

机构 * Cooperative Medianet Innovation Center, Shanghai Jiao Tong University(上海交通大学合作中位创新中心) School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19043 2025-07-16 cs.CV 79%

Self-Supervised Cross-Modal Text-Image Time Series Retrieval in Remote Sensing

Genc Hoxha, Olivér Angyal, Begüm Demir

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09174 2025-07-15 cs.CL 79%

RAMA: Retrieval-Augmented Multi-Agent Framework for Misinformation Detection in Multimodal Fact-Checking

Shuo Yang, Zijian Yu, Zhenzhe Ying, Yuqin Dai, Guoqing Wang, Jun Lan, Jinfeng Xu, Jinze Li, Edith C. H. Ngai

机构 * The University of Hong Kong(香港大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16035 2025-07-15 cs.LG cs.AI cs.IR 79%

Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding

Vishesh Tripathi, Tanmay Odapally, Indraneel Das, Uday Allu, Biddwan Ahmed

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments 11 pages, 1 Figure, 1 Table

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02357 2025-07-04 cs.CL 79%

Coling-UniA at SciVQA 2025: Few-Shot Example Retrieval and Confidence-Informed Ensembling for Multimodal Large Language Models

Christian Jaumann, Annemarie Friedrich, Rainer Lienhart

机构 * University of Augsburg(奥格斯堡大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at 5th Workshop on Scholarly Document Processing @ ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01066 2025-07-03 cs.IR cs.CV cs.LG 79%

Embedding-based Retrieval in Multimodal Content Moderation

Hanzhong Liang, Jinghao Shi, Xiang Shen, Zixuan Wang, Vera Wen, Ardalan Mehrani, Zhiqian Chen, Yifan Wu, Zhixin Zhang

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.CV

Comments Camera ready for SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01644 2025-07-01 cs.CL cs.LG 79%

Multimodal Contrastive Representation Learning in Augmented Biomedical Knowledge Graphs

Tien Dang, Viet Thanh Duy Nguyen, Minh Tuan Le, Truong-Son Hy

机构 * University of Alabama at Birmingham(阿拉巴马大学伯明翰分校) Washington University in St. Louis(华盛顿大学圣路易斯分校)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22056 2025-06-30 cs.AI 79%

Universal Retrieval for Multimodal Trajectory Modeling

Xuan Zhang, Ziyan Jiang, Rui Meng, Yifei Leng, Zhenbang Xiao, Zora Zhiruo Wang, Yanyi Shang, Dehan Kong

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments 18 pages, 3 figures, accepted by Workshop on Computer-use Agents @ ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21056 2025-06-27 cs.CV 79%

SAMURAI: Shape-Aware Multimodal Retrieval for 3D Object Identification

Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science, VNU-HCM(科学大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20070 2025-06-26 cs.IR cs.LG cs.MM 79%

Multimodal Information Retrieval for Open World with Edit Distance Weak Supervision

KMA Solaiman, Bharat Bhargava

机构 * Department of Computer Science Purdue University West Lafayette, IN 47906, USA(计算机科学系 帕克森大学 西拉法叶市,印第安纳州 47906,美国)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

Comments Submitted to ICDE'24. An earlier version of this paper appeared on TechRxiv: https://www.techrxiv.org/doi/full/10.36227/techrxiv.21990284.v1, uploaded on February 05, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏