arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3484 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3484 篇

2510.09894 2026-03-17 cs.AI cs.CY cs.LG 70%

Beyond AlphaEarth: Toward Human-Centered Geospatial Foundation Models via POI-Guided Contrastive Learning

超越AlphaEarth:通过POI引导对比学习实现以人为中心的地理空间基础模型

Junyuan Liu, Quan Qin, Guangsheng Dong, Xinglei Wang, Jiazhuang Feng, Zichao Zeng, Tao Cheng

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出AETHER框架,通过POI引导的多模态对齐,将AlphaEarth与以人为中心的城市分析结合,提升地理空间表示的可解释性与语言可访问性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06704 2026-03-10 cs.CV cs.LG 70%

On the Generalization Capacities of MLLMs for Spatial Intelligence

关于多模态大语言模型在空间智能中的泛化能力

Gongjie Zhang, Wenhao Li, Quanhao Qian, Jiuniu Wang, Deli Zhao, Shijian Lu, Ran Xu

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) HuPan Lab(虎朋实验室) Nanyang Technological University(南洋理工大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出Camera-Aware MLLM框架,通过注入相机内参、数据增强和蒸馏几何先验,提升多模态大语言模型在空间任务中的泛化能力与鲁棒性。

Comments ICLR 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04614 2026-03-06 cs.CV 70%

SGR3 Model: Scene Graph Retrieval-Reasoning Model in 3D

SGR3模型:三维场景图检索-推理模型

Zirui Wang, Ruiping Liu, Yufan Chen, Junwei Zheng, Weijia Fan, Kunyu Peng, Di Wen, Jiale Wei, Jiaming Zhang, Rainer Stiefelhagen

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Shenzhen University(深圳大学) Hunan University(湖南大学)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 SGR3模型通过多模态大语言模型与检索增强生成技术,实现无需训练的三维场景图生成,提升场景关系推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03617 2026-03-05 cs.CV 70%

RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation

RAGTrack: 基于检索增强生成的语言感知RGBT跟踪

Hao Li, Yuhao Wang, Wenning Hao, Pingping Zhang, Dong Wang, Huchuan Lu

机构 * College of Command and Control Engineering, Army Engineering University of PLA(中国人民解放军陆军工程大学指挥控制工程学院) School of Future Technology, Dalian University of Technology(大连理工大学未来技术学院) School of Information and Communication Engineering, Dalian University of Technology(大连理工大学信息与通信工程学院)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 RAGTrack通过引入语言指导的检索增强生成框架,解决RGBT跟踪中外观变化和模态间隙问题,实现稳健的目标定位。

Comments This work is accepted by CVPR2026. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01632 2026-03-03 cs.LG cs.AI 70%

DeLo: Dual Decomposed Low-Rank Experts Collaboration for Continual Missing Modality Learning

DeLo:双分解低秩专家协作用于持续缺失模态学习

Xiwei Liu, Yulong Li, Feilong Tang, Imran Razzak

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 DeLo通过双分解低秩专家架构解决持续缺失模态学习中的模态干扰问题,结合跨模态引导路由和任务键记忆实现高效推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04932 2026-03-03 cs.CV 70%

UniView: Enhancing Novel View Synthesis From A Single Image By Unifying Reference Features

UniView: 通过统一参考特征从单张图像增强新颖视角合成

Haowang Cui, Rui Chen, Jiaze Wang, Tao Guo, Zheng Qin

机构 * Tianjin University(天津大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 UniView通过统一参考特征提升单张图像新颖视角合成性能,采用检索增强系统和多模态大语言模型选择参考图像,并结合解耦三重注意力机制提高合成效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20530 2026-02-25 cs.LG cs.SD eess.AS 70%

Memory-guided Prototypical Co-occurrence Learning for Mixed Emotion Recognition

基于记忆的原型共现学习用于混合情绪识别

Ming Li, Yong-Jin Liu, Fang Liu, Huankun Sheng, Yeying Fan, Yixiang Wei, Minnan Luo, Weizhan Zhang, Wenping Wang

机构 * MOE-Key Laboratory of Pervasive Computing, Department of Computer Science and Technology, Tsinghua University(MOE-Key pervasive computing laboratory,计算机科学与技术系,清华大学) Key State Laboratory of Media Convergence and Communication, Communication University of China(媒体融合与传播关键实验室,中国传媒大学) Department of Computer Science and Computer Engineering, Texas A&M University(计算机科学与工程系,德克萨斯大学)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 eess.AS

AI总结 本文提出基于记忆的原型共现学习框架,通过多模态信号融合和原型关系蒸馏,提升混合情绪识别的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03098 2026-02-24 cs.LG cs.AI 70%

TextME: Bridging Unseen Modalities Through Text Descriptions

TextME: 通过文本描述弥合未见模态

Soyeon Hong, Jinchan Kim, Jaegook You, Seungtaek Choi, Suha Kwak, Hyunsouk Cho

机构 * Department of Artificial Intelligence, Ajou University, Suwon, South Korea Division of Language \& AI, Hankuk University of Foreign Studies, Seoul, Korea Graduate School of AI, POSTECH, Pohang, Korea Department of Software, Ajou University, Suwon, South Korea

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 TextME通过仅使用文本描述实现跨模态扩展,无需配对监督,有效弥合不同模态间的差距。

Comments Code available at https://github.com/SoyeonHH/TextME

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17402 2026-02-20 cs.AI 70%

A Contrastive Variational AutoEncoder for NSCLC Survival Prediction with Missing Modalities

用于NSCLC生存预测的对比变分自编码器,具有缺失模态

Michele Zanitti, Vanja Miskovic, Francesco Trovò, Alessandra Laura Giulia Pedrocchi, Ming Shen, Yan Kyaw Tun, Arsela Prelaj, Sokol Kosta

机构 * Department of Electronic Systems, Aalborg University, Copenhagen, Denmark(电子系统系,奥胡斯大学) Department of Electronics, Information and Bioengineering, Politecnico di Milano, Milan, Italy(电子、信息与生物工程系,米兰理工学院) Department of Medical Oncology, Istituto Nazionale dei Tumori, Milan, Italy(医学肿瘤学系,国家肿瘤研究所)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出了一种多模态对比变分自编码器,用于NSCLC生存预测,通过整合多种数据模态并处理缺失数据,提高预测的鲁棒性和准确性。

Comments Accepted at The 13th IEEE International Conference on Big Data (IEEE BigData 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21033 2026-02-03 cs.SD cs.AI 70%

SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization

SupCLAP:通过支持向量正则化控制音频-文本对比学习中的优化轨迹漂移

Jiehui Luo, Yuguo Yin, Yuxin Xie, Jinghan Ru, Xianwei Zhuang, Minghua He, Aofan Liu, Zihan Xiong, Dongchao Yang

机构 * Peking University(北京大学) Central Conservatory of Music(中央音乐学院) The Chinese University of Hong Kong(香港中文大学) University of Electronic Science and Technology of China(电子科技大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 SupCLAP通过支持向量正则化有效控制音频-文本对比学习中的优化轨迹漂移,提升多模态学习的稳定性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00131 2026-02-03 cs.CV cs.RO 70%

PovNet+: A Deep Learning Architecture for Socially Assistive Robots to Learn and Assist with Multiple Activities of Daily Living

PovNet+: 一种深度学习架构用于社交辅助机器人学习和协助多种日常活动

Fraser Robinson, Souren Pashangpour, Matthew Lisondra, Goldie Nejat

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 PovNet+是一种多模态深度学习架构,用于社交辅助机器人识别多种日常活动并主动发起辅助行为,提升了ADL分类准确率和人机交互能力。

Comments Submitted to Advanced Robotics (Taylor & Francis)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14259 2026-01-21 cs.CV 70%

ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation

ManipShield: 一种用于图像篡改检测、定位和解释的统一框架

Zitong Xu, Huiyu Duan, Xiaoyu Wang, Zhaolin Cai, Kaiwei Zhang, Qiang Hu, Jing Liu, Xiongkuo Min, Guangtao Zhai

机构 * Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University(上海交通大学图像通信与网络工程研究所) University of Electronic and Science Technology of China(电子科技大学) Tianjin University(天津大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 ManipShield基于多模态大语言模型,通过对比学习LoRA微调和任务特定解码器,实现图像篡改的统一检测、定位和解释。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13072 2025-12-16 cs.CV 70%

Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models

锻造动态记忆:基于检索的持续学习用于通用医学基础模型

Zizhi Chen, Yizhen Gao, Minghao Han, Yizhou Liu, Zhaoyu Chen, Dingkang Yang, Lihua Zhang

机构 * College of Intelligent Robotics and Advanced Manufacturing(智能机器人与先进制造学院) Fudan University(复旦大学) Fysics Intelligence Technologies Co., Ltd. (Fysics AI)(Fysics智能技术有限公司(Fysics AI)) School of Computer Science and Engineering(计算机科学与工程学院) Central South University(中南大学)

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 本文提出基于检索的持续学习方法,通过动态知识蒸馏和RAG技术,解决多模态医学模型在领域迁移和细粒度特征保留中的核心难题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18298 2025-11-25 cs.AI 70%

Cross-Disciplinary Knowledge Retrieval and Synthesis: A Compound AI Architecture for Scientific Discovery

跨学科知识检索与综合:一种用于科学发现的复合AI架构

Svitlana Volkova, Peter Bautista, Avinash Hiriyanna, Gabriel Ganberg, Isabel Erickson, Zachary Klinefelter, Nick Abele, Hsien-Te Kao, Grant Engberson

机构 * Aptima, Inc.(Aptima公司)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 BioSage通过整合LLMs与RAG,利用专门代理实现跨学科知识检索与综合,提升科学发现效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15308 2025-11-20 cs.CV 70%

Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language

Yan Xia, Letian Shi, Yilin Di, Joao F. Henriques, Daniel Cremers

机构 * School of Artificial Intelligence and Data Science, University of Science and Technology of China(人工智能与数据科学学院,中国科学技术大学) Technical University of Munich(慕尼黑技术大学) Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments This paper builds upon and extends our earlier conference paper Text2Loc presented at CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10390 2025-11-18 cs.CV 70%

HIBMatch: Hypergraph Information Bottleneck for Semi-supervised Alzheimer's Progression

Zhongying Deng, Shujun Wang, Angelica I Aviles-Rivero, Zoe Kourtzi, Carola-Bibiane Schönlieb

机构 * Department of Applied Mathematics and Theoretical Physics, University of Cambridge(应用数学与理论物理系,剑桥大学) Department of Biomedical Engineering, The Hong Kong Polytechnic University(生物医学工程系,香港理工大学) Research Institute for Artificial Intelligence of Things, The Hong Kong Polytechnic University(物联网人工智能研究所,香港理工大学) Yau Mathematical Sciences Centre, Tsinghua University(叶德平数学科学中心,清华大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to the IEEE Journal of Biomedical and Health Informatics (To appear)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05630 2025-11-11 q-bio.NC cs.AI 70%

BrainCSD: A Hierarchical Consistency-Driven MoE Foundation Model for Unified Connectome Synthesis and Multitask Brain Trait Prediction

Xiongri Shen, Jiaqi Wang, Yi Zhong, Zhenxi Song, Leilei Zhao, Liling Li, Yichen Wei, Lingyan Liang, Shuqiang Wang, Baiying Lei, Demao Deng, Zhiguo Zhang

机构 * Department of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术系) School of Intelligence Science and Engineering, College of Artificial Intelligence, Harbin Institute of Technology(哈尔滨工业大学智能科学与工程学院) School of Biomedical Engineering, National-Regional Key Technology Engineering Laboratory for Medical Ultrasound, Guangdong Key Laboratory for Biomedical, Measurements and Ultrasound Imaging, Shenzhen University Medical School, Shenzhen University(深圳大学医学院生物医学工程学院) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15252 2025-11-10 cs.CR cs.CL cs.IR 70%

Retrieval-Augmented Review Generation for Poisoning Recommender Systems

Shiyi Yang, Xinshu Li, Guanglin Zhou, Chen Wang, Xiwei Xu, Liming Zhu, Lina Yao

机构 * University of New South Wales and CSIRO’s Data61(新南威尔士大学和CSIRO的Data61)

专题命中 跨模态检索 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18623 2025-10-17 cs.CV 70%

Training-Free Personalization via Retrieval and Reasoning on Fingerprints

Deepayan Das, Davide Talon, Yiming Wang, Massimiliano Mancini, Elisa Ricci

机构 * University of Trento(特伦托大学) Fondazione Bruno Kessler(布鲁诺·凯瑟实验室)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12323 2025-10-15 cs.AI 70%

RAG-Anything: All-in-One RAG Framework

Zirui Guo, Xubin Ren, Lingrui Xu, Jiahao Zhang, Chao Huang

机构 * The University of Hong Kong(香港大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21524 2025-10-14 cs.CV cs.LG stat.ML 70%

Learning Shared Representations from Unpaired Data

Amitai Yacobi, Nir Ben-Ari, Ronen Talmon, Uri Shaham

机构 * Department of Computer Science Bar-Ilan University(巴伊兰大学计算机科学系) Electrical and Computer Engineering Technion(技术学院电子与计算机工程系)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02790 2025-10-06 cs.CV cs.AI cs.CL cs.MM 70%

MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding

Jingyuan Deng, Yujiu Yang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments accepted to emnlp2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26330 2025-10-01 cs.CV cs.IR 70%

SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval

Ren-Di Wu, Yu-Yen Lin, Huei-Fang Yang

机构 * National Sun Yat-sen University(国立中山大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 20 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13146 2025-09-23 cs.CV cs.LG 70%

Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization

Shuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai, Yueqi Wang, Chan-Wei Hu, Chengxuan Qian, Huaxiu Yao, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯大学) University of Michigan(密歇根大学) UIUC(伊利诺伊大学香槟分校) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Published at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14746 2025-09-19 cs.CV cs.IR 70%

Chain-of-Thought Re-ranking for Image Retrieval Tasks

Shangrong Wu, Yanghong Zhou, Yang Chen, Feng Zhang, P. Y. Mok

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12994 2025-09-17 cs.CL 70%

SitLLM: Large Language Models for Sitting Posture Health Understanding via Pressure Sensor Data

Jian Gao, Fufangchen Zhao, Yiyang Zhang, Danfeng Yan

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09118 2025-09-12 cs.CV 70%

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

Tianlu Zheng, Yifan Zhang, Xiang An, Ziyong Feng, Kaicheng Yang, Qichuan Ding

机构 * Northeastern University(东北大学) South China University of Technology(南方科技大学) DeepGlint

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by EMNLP2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04376 2025-09-08 cs.CV 70%

AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search

Hao Ju, Hu Zhang, Zhedong Zheng

机构 * Faculty of Science and Technology and Institute of Collaborative Innovation, University of Macau(科技学院和协同创新研究所,澳门大学) CSIRO Data61(CSIRO数据61)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21539 2025-09-01 cs.CV 70%

HCCM: Hierarchical Cross-Granularity Contrastive and Matching Learning for Natural Language-Guided Drones

Hao Ruan, Jinliang Lin, Yingxin Lai, Zhiming Luo, Shaozi Li

机构 * Department of Artificial Intelligence, Xiamen University(人工智能学院,厦门大学)

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted by ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06104 2025-08-11 cs.CV 70%

MCA: 2D-3D Retrieval with Noisy Labels via Multi-level Adaptive Correction and Alignment

Gui Zou, Chaofan Gan, Chern Hong Lim, Supavadee Aramvith, Weiyao Lin

机构 * Shanghai Jiao Tong University, China(上海交通大学) Monash University, Malaysia(墨尔本大学) Chulalongkorn University, Thailand(朱拉隆功大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments ICMEW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏