arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3496 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3496 篇

2412.16919 2025-08-12 cs.CV 57%

TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction

Xuying Zhang, Yutong Liu, Yangguang Li, Renrui Zhang, Yufei Liu, Kai Wang, Wanli Ouyang, Zhiwei Xiong, Peng Gao, Qibin Hou, Ming-Ming Cheng

机构 * VCIP, CS, Nankai University(南开大学计算机科学与技术学院) NKIARI, Shenzhen Futian(深圳未来科技研究院) USTC(University of Science and Technology of China) CUHK MMLab(香港中文大学MMLab) VAST(中国科学院自动化研究所) Shanghai AI Lab(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted at ICCV 2025. Project page: https://github.com/HVision-NKU/TAR3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14428 2025-08-11 cs.CV cs.LG q-bio.QM 57%

WildSAT: Learning Satellite Image Representations from Wildlife Observations

Rangel Daroya, Elijah Cole, Oisin Mac Aodha, Grant Van Horn, Subhransu Maji

机构 * University of Massachusetts, Amherst(马萨诸塞大学阿默斯特分校) GenBio AI University of Edinburgh(爱丁堡大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23029 2025-08-08 cs.CL 57%

Uncovering Visual-Semantic Psycholinguistic Properties from the Distributional Structure of Text Embedding Space

Si Wu, Sebastian Bruch

机构 * Northeastern University(东北大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CL

Comments The camera-ready version for ACL 2025 in Vienna

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23736 2025-08-07 cs.CV cs.IR 57%

Modality and Task Adaptation for Enhanced Zero-shot Composed Image Retrieval

Haiwen Li, Fei Su, Zhicheng Zhao

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23735 2025-08-05 cs.RO cs.AI cs.MA 57%

Distributed AI Agents for Cognitive Underwater Robot Autonomy

Markus Buchholz, Ignacio Carlucho, Michele Grimaldi, Yvan R. Petillot

机构 * School of Engineering & Physical Sciences, Heriot-Watt University(工程与物理科学学院,赫里奥特-瓦特大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00217 2025-08-04 cs.CL cs.DB cs.LG 57%

Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges

Xiaofeng Wu, Alan Ritter, Wei Xu

机构 * College of Computing, Georgia Institute of Technology(计算学院、佐治亚理工学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23217 2025-08-01 cs.LG cs.AI 57%

Zero-Shot Document Understanding using Pseudo Table of Contents-Guided Retrieval-Augmented Generation

Hyeon Seong Jeong, Sangwoo Jo, Byeong Hyun Yoon, Yoonseok Heo, Haedong Jeong, Taehoon Kim

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21015 2025-07-29 cs.CV 57%

Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions

Licai Sun, Xingxun Jiang, Haoyu Chen, Yante Li, Zheng Lian, Biu Liu, Yuan Zong, Wenming Zheng, Jukka M. Leppänen, Guoying Zhao

机构 * University of Oulu(奥卢大学) Southeast University(东南大学) University of Turku(图尔库大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18881 2025-07-28 cs.CV cs.RO 57%

Perspective from a Higher Dimension: Can 3D Geometric Priors Help Visual Floorplan Localization?

Bolei Chen, Jiaxu Kang, Haonan Yang, Ping Zhong, Jianxin Wang

机构 * School of Computer Science Engineering, Central South University Changsha Hunan China Engineering, Central South University

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12591 2025-07-28 cs.CL 57%

LLMs are Also Effective Embedding Models: An In-depth Overview

Chongyang Tao, Tao Shen, Shen Gao, Junshuo Zhang, Zhen Li, Kai Hua, Wenpeng Hu, Zhengwei Tao, Shuai Ma

机构 * Beihang University(北航大学) SKLSDE Lab, Beihang University(北航SKLSDE实验室) University of Technology Sydney(悉尼大学) University of Electronic Science and Technology of China(电子科技大学) Peking University(北京大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL

Comments 38 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20292 2025-07-25 cs.CV cs.LG 57%

Visual Adaptive Prompting for Compositional Zero-Shot Learning

Kyle Stein, Arash Mahyari, Guillermo Francia, Eman El-Sheikh

机构 * University of West Florida(乌斯托尔大学) Florida Institute for Human and Machine Cognition (IHMC)(佛罗里达人类与机器认知研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17183 2025-07-21 eess.IV cs.CV 57%

Large-Vocabulary Segmentation for Medical Images with Text Prompts

Ziheng Zhao, Yao Zhang, Chaoyi Wu, Xiaoman Zhang, Xiao Zhou, Ya Zhang, Yanfeng Wang, Weidi Xie

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 74 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01048 2025-07-17 cs.CV cs.CR cs.LG 57%

How does Watermarking Affect Visual Language Models in Document Understanding?

Chunxue Xu, Yiwei Wang, Bryan Hooi, Yujun Cai, Songze Li

机构 * Southeast University, China(东南大学) University of California, Merced, USA(加州大学默塞德分校) National University of Singapore, Singapore(新加坡国立大学) The University of Queensland, Australia(昆士兰大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11479 2025-07-16 cs.AI cs.GR cs.HC 57%

Perspective-Aware AI in Extended Reality

Daniel Platnick, Matti Gruener, Marjan Alirezaie, Kent Larson, Dava J. Newman, Hossein Rahnama

机构 * Flybits Labs(Flybits实验室) Creative Ai Hub(创意人工智能中心) Toronto Metropolitan University(多伦多 Metropolitan 大学) MIT Media Lab(MIT媒体实验室)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments Accepted to the International Conference on eXtended Reality (2025), 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06979 2025-07-10 cs.LG cs.CV 57%

A Principled Framework for Multi-View Contrastive Learning

Panagiotis Koromilas, Efthymios Georgiou, Giorgos Bouritsas, Theodoros Giannakopoulos, Mihalis A. Nicolaou, Yannis Panagakis

机构 * Department of Informatics and Telecommunications, National and Kapodistrian University of Athens(信息与通信技术系,雅典国家与卡波迪斯特里亚大学) Archimedes AI/Athena Research Center(阿基米德AI/阿塔纳研究中心) ILSP/Athena Research Center(ILSP/阿塔纳研究中心) The Cyprus Institute(塞浦路斯研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00160 2025-07-08 cs.CL 57%

Emergency Department Decision Support using Clinical Pseudo-notes

Simon A. Lee, Sujay Jain, Alex Chen, Kyoka Ono, Jennifer Fang, Akos Rudas, Jeffrey N. Chiang

机构 * Department of Computational Medicine, University of California, Los Angeles, CA 90095 USA(计算医学系,加州大学洛杉矶分校) Department of Neurosurgery, University of California, Los Angeles, CA 90095 USA(神经外科系,加州大学洛杉矶分校) Department of Electrical and Computer Engineering, University of California at Los Angeles, Los Angeles, CA 90095 USA(电气与计算机工程系,加州大学洛杉矶分校) Harbor-UCLA Medical Center, Department of Emergency Medicine, Torrance, CA(Harbor-UCLA医疗中心,急诊医学部) University of California, Los Angeles, Department of Emergency Medicine, Los Angeles, California(加州大学洛杉矶分校,急诊医学部) Department of Statistics and Data Science University of California, Los Angeles(统计与数据科学系,加州大学洛杉矶分校) Department of Natural Sciences International Christian University, Mitaka, Tokyo, Japan(自然科学系,国际基督教大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Journal ref npj Digital Medicine 8 (1), 394, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03250 2025-07-08 cs.CV cs.LG 57%

Subject Invariant Contrastive Learning for Human Activity Recognition

Yavuz Yarici, Kiran Kokilepersaud, Mohit Prabhushankar, Ghassan AlRegib

机构 * Georgia Institute of Technology(佐治亚理工学院) Center for Signal and Information Processing CSIP(信号与信息处理中心)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11281 2025-07-02 cs.CV q-bio.QM 57%

DynaCLR: Contrastive Learning of Cellular Dynamics with Temporal Regularization

Eduardo Hirata-Miyasaki, Soorya Pradeep, Ziwen Liu, Alishba Imran, Taylla Milena Theodoro, Ivan E. Ivanov, Sudip Khadka, See-Chi Lee, Michelle Grunberg, Hunter Woosley, Madhura Bhave, Carolina Arias, Shalin B. Mehta

机构 * Chan Zuckerberg Biohub(查纳·泽伯生物枢纽) University of California Berkeley(加州大学伯克利分校)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments 30 pages, 6 figures, 13 appendix figures, 5 videos (ancillary files)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23827 2025-07-01 cs.CV 57%

Spatially Gene Expression Prediction using Dual-Scale Contrastive Learning

Mingcheng Qu, Yuncong Wu, Donglin Di, Yue Gao, Tonghua Su, Yang Song, Lei Fan

机构 * Faculty of Computing, Harbin Institute of Technology, Harbin, China(哈尔滨工业大学计算机学院) School of Astronautics, Harbin Institute of Technology, Harbin, China(哈尔滨工业大学航天学院) School of Software, Tsinghua University, Beijing, China(清华大学软件学院) School of Computer Science and Engineering, UNSW, Sydney, Australia(新南威尔士大学计算机科学与工程学院)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

Comments Our paper has been accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12494 2025-07-01 cs.CL cs.IR 57%

FlexRAG: A Flexible and Comprehensive Framework for Retrieval-Augmented Generation

Zhuocheng Zhang, Yang Feng, Min Zhang

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS)(中国科学院智能信息处理重点实验室) University of Chinese Academy of Sciences, China(中国科学院大学) Key Laboratory of AI Safety, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室) Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen), China(哈尔滨工业大学(深圳)计算机与智能信息研究院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

Comments Accepted by ACL 2025 Demo

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11134 2025-07-01 cs.CV 57%

Visual Re-Ranking with Non-Visual Side Information

Gustav Hanning, Gabrielle Flood, Viktor Larsson

机构 * Lund University(吕勒欧大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted at Scandinavian Conference on Image Analysis (SCIA) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21934 2025-06-30 cs.IR cs.CV 57%

CAL-RAG: Retrieval-Augmented Multi-Agent Generation for Content-Aware Layout Design

Najmeh Forouzandehmehr, Reza Yousefi Maragheh, Sriram Kollipara, Kai Zhao, Topojoy Biswas, Evren Korpeoglu, Kannan Achan

机构 * Walmart Global Tech(沃尔玛全球技术)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13897 2025-06-27 cs.CV 57%

DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity Understanding

Thomas Kreutz, Max Mühlhäuser, Alejandro Sanchez Guinea

机构 * Telekooperation Lab, Technical University Darmstadt(德累斯顿技术大学电信协作实验室) NTT DATA, Luxembourg(NTT DATA卢森堡)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20612 2025-06-27 cs.LG cs.CV 57%

Discovering Global False Negatives On the Fly for Self-supervised Contrastive Learning

Vicente Balmaseda, Bokun Wang, Ching-Long Lin, Tianbao Yang

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16065 2025-06-26 cs.IR cs.CL 57%

Aug2Search: Enhancing Facebook Marketplace Search with LLM-Generated Synthetic Data Augmentation

Ruijie Xi, He Ba, Hao Yuan, Rishu Agrawal, Yuxin Tian, Ruoyan Kong, Arul Prakash

机构 * North Carolina State University(北卡罗来纳州立大学) Meta

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17886 2025-06-25 cs.SD eess.AS 57%

GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models

Julien Guinot, Elio Quinton, György Fazekas

专题命中 跨模态检索 :multimodal(abstract);分类 eess.AS

Comments Accepted to ISMIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18856 2025-06-24 cs.CV 57%

RAG-6DPose: Retrieval-Augmented 6D Pose Estimation via Leveraging CAD as Knowledge Base

Kuanning Wang, Yuqian Fu, Tianyu Wang, Yanwei Fu, Longfei Liang, Yu-Gang Jiang, Xiangyang Xue

机构 * Fudan University(复旦大学) NeuhHelium Co.,Ltd.(NeuhHelium公司)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15802 2025-06-24 cs.CV 57%

Visual Prompt Engineering for Vision Language Models in Radiology

Stefan Denner, Markus Bujotzek, Dimitrios Bounias, David Zimmerer, Raphael Stock, Klaus Maier-Hein

机构 * Division of Medical Image Computing, German Cancer Research Center, Heidelberg, Germany(德国癌症研究中心医学图像计算部) Faculty of Mathematics and Computer Science, Heidelberg University, Heidelberg, Germany(海德堡大学数学与计算机科学学院) Medical Faculty Heidelberg, University of Heidelberg, Heidelberg, Germany(海德堡大学医学学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments Accepted at ECCV 2024 Workshop on Emergent Visual Abilities and Limits of Foundation Models & Medical Imaging with Deep Learning 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16683 2025-06-23 cs.IR cs.AI 57%

A Simple Contrastive Framework Of Item Tokenization For Generative Recommendation

Penglong Zhai, Yifang Yuan, Fanyi Di, Jie Li, Yue Liu, Chen Li, Jie Huang, Sicong Wang, Yao Xu, Xin Li

机构 * AMAP, Alibaba Group(阿里集团AMAP)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

Comments 12 pages,7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13496 2025-06-23 cs.CV cs.IR cs.LG 57%

Hierarchical Multi-Positive Contrastive Learning for Patent Image Retrieval

Kshitij Kavimandan, Angelos Nalmpantis, Emma Beauxis-Aussalet, Robert-Jan Sips

机构 * Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) TKH AI(TKH人工智能)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 5 pages, 3 figures, Accepted as a short paper at the 6th Workshop on Patent Text Mining and Semantic Technologies (PatentSemTech 2025), co-located with SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏