arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3484 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3484 篇

2507.09459 2025-07-15 cs.CV cs.RO 70%

SegVec3D: A Method for Vector Embedding of 3D Objects Oriented Towards Robot manipulation

Zhihan Kang, Boyu Wang

机构 * Northwestern Polytechnical University(西北工业大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Undergraduate Theis; 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07902 2025-07-11 cs.CV 70%

MIRA: A Novel Framework for Fusing Modalities in Medical RAG

Jinhong Wang, Tajamul Ashraf, Zongyan Han, Jorma Laaksonen, Rao Mohammad Anwer

机构 * Department of Computer Vision, MBZUAI(视觉计算系,MBZUAI) Department of Computer Science, Aalto University(计算机科学系,阿alto大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04735 2025-07-08 cs.CV 70%

An analysis of vision-language models for fabric retrieval

Francesco Giuliari, Asif Khan Pattan, Mohamed Lamine Mekhalfi, Fabio Poiesi

机构 * Fondazione Bruno Kessler(布雷诺-科塞勒基金会)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at Ital-IA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13066 2025-06-17 cs.CL 70%

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design

Kai Lan, Jiayong Zhu, Jiangtong Li, Dawei Cheng, Guang Chen, Changjun Jiang

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

Comments 26 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11082 2025-06-10 cs.CV cs.AI cs.CL cs.MM 70%

Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning

Chen Jiang, Hong Liu, Xuzheng Yu, Qing Wang, Yuan Cheng, Jia Xu, Zhongyi Liu, Qingpei Guo, Wei Chu, Ming Yang, Yuan Qi

机构 * Artificial Intelligence Innovation and Incubation Institute, Fudan University(复旦大学人工智能创新与孵化院) Ant Group(蚂蚁集团)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06602 2025-06-10 cs.CV 70%

Zero Shot Composed Image Retrieval

Santhosh Kakarla, Gautama Shastry Bulusu Venkata

机构 * George Mason University(乔治·玛森大学)

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05443 2025-06-09 cs.LG cs.AI q-bio.GN 70%

UniPTMs: The First Unified Multi-type PTM Site Prediction Model via Master-Slave Architecture-Based Multi-Stage Fusion Strategy and Hierarchical Contrastive Loss

Yiyu Lin, Yan Wang, You Zhou, Xinye Ni, Jiahui Wu, Sen Yang

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17709 2025-06-06 cs.CV cs.AI cs.CL cs.LG cs.MM 70%

Contrastive Visual Data Augmentation

Yu Zhou, Bingxuan Li, Mohan Tang, Xiaomeng Jin, Te-Lin Wu, Kuan-Hao Huang, Heng Ji, Kai-Wei Chang, Nanyun Peng

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Journal ref ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03148 2025-06-04 cs.CV 70%

Self-Supervised Spatial Correspondence Across Modalities

Ayush Shrivastava, Andrew Owens

机构 * University of Michigan(密歇根大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments CVPR 2025. Project link: https://www.ayshrv.com/cmrw . Code: https://github.com/ayshrv/cmrw

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02291 2025-06-04 cs.CV cs.IR 70%

Entity Image and Mixed-Modal Image Retrieval Datasets

Cristian-Ioan Blaga, Paul Suganthan, Sahil Dua, Krishna Srinivasan, Enrique Alfonseca, Peter Dornbach, Tom Duerig, Imed Zitouni, Zhe Dong

机构 * Google Switzerland(谷歌瑞士分公司) Microsoft AI(微软人工智能)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24441 2025-06-02 cs.CV 70%

SORCE: Small Object Retrieval in Complex Environments

Chunxu Liu, Chi Xie, Xiaxu Chen, Wei Li, Feng Zhu, Rui Zhao, Limin Wang

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Sensetime Research(商汤科技研究院) Tongji University(同济大学) Beijing Institute of Technology(北京理工大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Project Page: https://github.com/MCG-NJU/SORCE

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02131 2025-05-08 cs.LG cs.CL 70%

Boosting Masked ECG-Text Auto-Encoders as Discriminative Learners

Hung Manh Pham, Aaqib Saeed, Dong Ma

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16352 2025-04-24 cs.IR cs.AI 70%

Disentangling and Generating Modalities for Recommendation in Missing Modality Scenarios

Jiwan Kim, Hongseok Kang, Sein Kim, Kibum Kim, Chanyoung Park

机构 * KAIST(韩国科学技术院)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

Comments SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09518 2025-04-15 cs.CV 70%

3D CoCa: Contrastive Learners are 3D Captioners

Ting Huang, Zeyu Zhang, Yemin Wang, Hao Tang

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08342 2025-03-13 cs.CV 70%

Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs

Chongjun Tu, Peng Ye, Dongzhan Zhou, Lei Bai, Gang Yu, Tao Chen, Wanli Ouyang

专题命中 跨模态检索 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18733 2025-02-27 cs.LG cs.AI 70%

Cross-Modality Investigation on WESAD Stress Classification

Eric Oliver, Sagnik Dakshit

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08438 2025-02-13 cs.CV cs.AI cs.CL cs.IR cs.MM 70%

Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions

Prajwal Gatti, Kshitij Parikh, Dhriti Prasanna Paul, Manish Gupta, Anand Mishra

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at AAAI 2024, 9 pages. Project Website: https://vl2g.github.io/projects/cstbir

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10501 2025-02-10 cs.CV 70%

Enhancing medical vision-language contrastive learning via inter-matching relation modelling

Mingjian Li, Mingyuan Meng, Michael Fulham, David Dagan Feng, Lei Bi, Jinman Kim

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Published at IEEE Transactions on Medical Imaging

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.07331 2025-02-04 cs.CV cs.LG 70%

Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering

Weixi Weng, Jieming Zhu, Xiaojun Meng, Hao Zhang, Rui Zhang, Chun Yuan

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14369 2025-01-27 cs.CV 70%

Low-rank Prompt Interaction for Continual Vision-Language Retrieval

Weicai Yan, Ye Wang, Wang Lin, Zirun Guo, Zhou Zhao, Tao Jin

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13277 2025-01-24 cs.CV 70%

MEDFORM: A Foundation Model for Contrastive Learning of CT Imaging and Clinical Numeric Data in Multi-Cancer Analysis

Daeun Jung, Jaehyeok Jang, Sooyoung Jang, Yu Rang Park

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 8 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13372 2024-12-19 cs.CV 70%

Restore Anything Model via Efficient Degradation Adaptation

Bin Ren, Eduard Zamfir, Zongwei Wu, Yawei Li, Yidi Li, Danda Pani Paudel, Radu Timofte, Ming-Hsuan Yang, Nicu Sebe

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Efficient Any Image Restoration

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11793 2024-12-16 cs.AI 70%

Leveraging Chemistry Foundation Models to Facilitate Structure Focused Retrieval Augmented Generation in Multi-Agent Workflows for Catalyst and Materials Design

Nathaniel H. Park, Tiffany J. Callahan, James L. Hedrick, Tim Erdmann, Sara Capponi

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08715 2024-12-12 cs.CV 70%

Retrieval Augmented Recipe Generation

Guoshan Liu, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yu-Gang Jiang

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

Comments ACCEPT on IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05756 2024-12-10 cs.CV 70%

Compositional Image Retrieval via Instruction-Aware Contrastive Learning

Wenliang Zhong, Weizhi An, Feng Jiang, Hehuan Ma, Yuzhi Guo, Junzhou Huang

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00436 2024-10-02 cs.RO cs.CV 70%

Task Success Prediction for Open-Vocabulary Manipulation Based on Multi-Level Aligned Representations

Miyu Goko, Motonari Kambara, Daichi Saito, Seitaro Otsuki, Komei Sugiura

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted for presentation at CoRL2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19439 2024-10-01 cs.CV 70%

Contrastive ground-level image and remote sensing pre-training improves representation learning for natural world imagery

Andy V. Huynh, Lauren E. Gillespie, Jael Lopez-Saucedo, Claire Tang, Rohan Sikand, Moisés Expósito-Alonso

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted to ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17779 2024-09-30 cs.CV 70%

DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction

Chaofan Gan, Yuanpeng Tu, Yuxi Li, Weiyao Lin

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15810 2024-09-25 cs.CV 70%

Hyperbolic Image-and-Pointcloud Contrastive Learning for 3D Classification

Naiwen Hu, Haozhe Cheng, Yifan Xie, Pengcheng Shi, Jihua Zhu

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at IROS2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11795 2024-09-11 cs.CV 70%

PoseScript: Linking 3D Human Poses and Natural Language

Ginger Delmas, Philippe Weinzaepfel, Thomas Lucas, Francesc Moreno-Noguer, Grégory Rogez

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments TPAMI 2024, extended version of the ECCV 2022 paper

详情

展开后加载摘要…

URL PDF HTML 收藏