arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3463 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3463 篇

2602.02004 2026-02-03 cs.CV cs.AI 81%

ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning

ClueTracer: 问题到视觉线索追踪用于无训练 hallucination 抑制在多模态推理

Gongli Xi, Kun Wang, Zeming Gao, Huahui Yi, Haolang Lu, Ye Tian, Wendong Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanyang Technological University(南洋理工大学) West China Biomedical Big Data Center(西京生物大数据中心)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 ClueTracer通过问题到视觉线索追踪,无训练抑制多模态推理中的幻觉,提升推理和非推理任务性能。

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01561 2026-02-03 cs.CV cs.AI 81%

Multimodal UNcommonsense: From Odd to Ordinary and Ordinary to Odd

多模态UNcommonsense:从奇特到普通和从普通到奇特

Yejin Son, Saejin Kim, Dongjun Min, Younjae Yu

机构 * Yonsei University(延世大学) Seoul National University(首尔国立大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 多模态UNcommonsense通过R-ICL框架提升模型在非典型场景下的推理能力,实现从奇特到普通和从普通到奇特的转换。

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19060 2026-01-28 cs.CV cs.AI 81%

Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models

基于像素的检索用于知识型大型多模态模型

Jeonghwan Kim, Renjie Tao, Sanat Sharma, Jiaqi Wang, Kai Sun, Zhaojiang Lin, Seungwhan Moon, Lambert Mathias, Anuj Kumar, Heng Ji, Xin Luna Dong

机构 * Meta Reality Labs(Meta现实实验室) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 PixSearch是一种端到端的多模态模型,通过像素级检索和统一感知与推理,提升视觉问答任务中的事实一致性与泛化能力。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02729 2026-01-26 cs.CL cs.AI cs.IR 81%

Unified Multimodal Interleaved Document Representation for Retrieval

统一多模态交错文档表示用于检索

Jaewoo Lee, Joonho Ko, Jinheon Baek, Soyeong Jeong, Sung Ju Hwang

机构 * University of North Carolina Chapel Hill(北卡罗来纳大学教堂山分校) KAIST(韩国科学技术院) DeepAuto

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出统一多模态交错文档表示方法,通过整合文本、图像和表格信息提升信息检索性能。

Comments EACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07780 2026-01-21 cs.CV cs.AI 81%

Semantic-Consistent Bidirectional Contrastive Hashing for Noisy Multi-Label Cross-Modal Retrieval

语义一致的双向对比哈希用于噪声多标签跨模态检索

Likang Peng, Chao Su, Wenyuan Wu, Yuan Sun, Dezhong Peng, Xi Peng, Xu Wang

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出语义一致的双向对比哈希方法,通过跨模态语义一致性分类和双向软对比哈希模块,有效应对多标签数据中的噪声问题,提升跨模态检索的鲁棒性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05600 2026-01-12 cs.CV cs.CL cs.LG 81%

SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes

SceneAlign: 在复杂视觉场景中将多模态推理对齐到场景图

Chuhan Wang, Xintong Li, Jennifer Yuntong Zhang, Junda Wu, Chengkai Huang, Lina Yao, Julian McAuley, Jingbo Shang

机构 * University of California, San Diego(加州大学圣地亚哥分校) University of Toronto(多伦多大学) University of New South Wales(新南威尔士大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 SceneAlign通过利用场景图进行结构干预,提升多模态推理在复杂视觉场景中的准确性和忠实性。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04571 2026-01-09 cs.AI cs.MM 81%

Enhancing Multimodal Retrieval via Complementary Information Extraction and Alignment

通过互补信息提取与对齐增强多模态检索

Delong Zeng, Yuexiang Xie, Yaliang Li, Ying Shen

机构 * School of Intelligent Systems Engineering, Sun Yat-sen University(1 智能系统工程学院,中山大学) Alibaba Group(2 阿里巴巴集团) Guangdong Provincial Key Laboratory of Fire Science and Intelligent Emergency Technology(3 广东省火灾科学与智能应急技术重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI、cs.MM

AI总结 CIEA通过互补信息提取与对齐提升多模态检索效果,实现对图像和文本统一潜在空间的建模,并在多个基准上取得显著优势。

Comments Accepted by ACL'2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04127 2026-01-08 cs.CV cs.AI 81%

Pixel-Wise Multimodal Contrastive Learning for Remote Sensing Images

像素级多模态对比学习用于遥感图像

Leandro Stival, Ricardo da Silva Torres, Helio Pedrini

机构 * Wageningen University & Research(瓦赫宁根大学与研究学院) Institute of Computing, University of Campinas (UNICAMP)(计算学院,坎皮纳斯大学(UNICAMP))

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出像素级多模态对比学习方法,通过二维表示和对比学习提升遥感图像中特征提取效果,优于现有方法。

Comments 21 pages, 9 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25270 2026-01-06 cs.LG cs.AI cs.CV 81%

InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interactions

InfMasking:通过对比多模态交互释放协同信息

Liangjian Wen, Qun Dai, Jianzhuang Liu, Jiangtao Zheng, Yong Dai, Dongkai Wang, Zhao Kang, Jun Wang, Zenglin Xu, Jiang Duan

机构 * School of Computing and Artificial Intelligence(计算与人工智能学院) Engineering Research Center of Intelligent Finance Ministry of Education(教育部长智能金融工程研究中心) Shenzhen Institutes of Advanced Technology Chinese Academy of Sciences(中国科学院深圳先进技术研究院) University of Electronic Science and Technology of China(电子科技大学) Shanghai Academy of AI for Science(上海人工智能科学研究院) Artificial Intelligence Innovation and Incubation Institute Fudan University(复旦大学人工智能创新与孵化院) Artificial Intelligence and Digital Finance Key Laboratory of Sichuan Province(四川省人工智能与数字金融重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 InfMasking通过无限遮蔽策略增强多模态协同信息提取,实现多模态表示学习中的协同信息优化与性能提升。

Comments Conference on Neural Information Processing Systems (NeurIPS) 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18987 2025-12-23 cs.RO cs.CL cs.CV 81%

Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation

语义可感知的多模态检索:基于具身记忆的层次化移动操作

Ryosuke Korekata, Quanting Xie, Yonatan Bisk, Komei Sugiura

机构 * Keio University(keio大学) Keio AI Research Center(keio人工智能研究中心) Carnegie Mellon University(卡内基梅隆大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本研究提出Affordance RAG框架,通过构建具有可操作性的具身记忆,提升机器人在开放词汇移动操作中的检索性能和任务成功率。

Comments Accepted to IEEE RA-L, with presentation at ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12935 2025-12-16 cs.CV cs.AI cs.IR 81%

Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion

通过级联嵌入-重排序和时间感知得分融合实现统一的交互式多模态时刻检索

Toan Le Ngo Thanh, Phat Ha Huu, Tan Nguyen Dang Duy, Thong Nguyen Le Minh, Anh Nguyen Nhu Tinh

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出通过级联嵌入-重排序和时间感知得分融合实现统一的多模态时刻检索系统,解决跨模态噪声、时间序列连贯性和模态选择问题。

Comments Accepted at AAAI Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02454 2025-12-16 cs.CL cs.AI 81%

Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports From Scratch with Agentic Framework

多模态深度研究员:从零开始生成文本-图表交织报告的代理框架

Zhaorui Yang, Bo Pan, Han Wang, Yiyao Wang, Xingyu Liu, Luoxuan Weng, Yingchaojie Feng, Haozhe Feng, Minfeng Zhu, Bo Zhang, Wei Chen

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 多模态深度研究员通过代理框架实现从零开始生成文本-图表交织报告,利用FDV结构化文本表示提升可视化生成质量。

Comments AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04741 2025-12-02 cs.AI cs.CL cs.HC 81%

Question Answering for Decisionmaking in Green Building Design: A Multimodal Data Reasoning Method Driven by Large Language Models

绿色建筑设计中的决策制定问答:一种由大语言模型驱动的多模态数据推理方法

Yihui Li, Xiaoyue Yan, Hao Zhou, Borong Lin

机构 * School of Architecture(建筑学院) Tsinghua University(清华大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本研究提出GreenQA,一种由大语言模型驱动的多模态数据推理方法,用于提升绿色建筑设计决策效率。

Comments Published at Association for Computer Aided Design in Architecture (ACADIA) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09878 2025-11-24 cs.CV cs.AI cs.RO 81%

CleverDistiller: Simple and Spatially Consistent Cross-modal Distillation

CleverDistiller: 简单且空间一致的跨模态知识蒸馏

Hariprasath Govindarajan, Maciej K. Wozniak, Marvin Klingner, Camille Maurice, B Ravi Kiran, Senthil Yogamani

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 CleverDistiller通过简单有效的跨模态知识蒸馏方法,在自动驾驶任务中实现语义分割和3D目标检测的高性能表现。

Comments Accepted to BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19579 2025-11-20 cs.CV cs.AI cs.LG 81%

Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration

Francisco Mena, Dino Ienco, Cassio F. Dantas, Roberto Interdonato, Andreas Dengel

机构 * Department of Computer Science, University of Kaiserslautern-Landau (RPTU)(科斯拉尔特伦大学计算机科学系) SDS, German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) INRAE, UMR TETIS, University of Montpellier(蒙彼利埃大学UMR TETIS) CIRAD, UMR TETIS, University of Montpellier(蒙彼利埃大学UMR TETIS) INRIA, EVERGREEN, University of Montpellier(蒙彼利埃大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted at the Machine Learning journal, CfP: Discovery Science 2024

Journal ref Machine Learning 114, 279 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11305 2025-11-19 cs.IR cs.AI cs.CV cs.LG 81%

MOON Embedding: Multimodal Representation Learning for E-commerce Search Advertising

Chenghan Fu, Daoze Zhang, Yukang Lin, Zhanheng Nie, Xiang Zhang, Jianyu Liu, Yueran Liu, Wanxian Guan, Pengjie Wang, Jian Xu, Bo Zheng

机构 * Alibaba Group(阿里巴巴集团) Taobao(淘宝)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 31 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05571 2025-11-11 cs.CV cs.AI 81%

C3-Diff: Super-resolving Spatial Transcriptomics via Cross-modal Cross-content Contrastive Diffusion Modelling

Xiaofei Wang, Stephen Price, Chao Li

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05404 2025-11-10 cs.CV cs.AI 81%

Multi-modal Loop Closure Detection with Foundation Models in Severely Unstructured Environments

Laura Alejandra Encinar Gonzalez, John Folkesson, Rudolph Triebel, Riccardo Giubilato

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments Under review for ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22900 2025-11-10 cs.CV cs.CL 81%

MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering

Mai A. Shaaban, Tausifa Jan Saleem, Vijay Ram Papineni, Mohammad Yaqub

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Department of Mathematics and Computer Science, Faculty of Science, Alexandria University(亚历山大大学数学与计算机科学系) Sheikh Shakhbout Medical City(谢赫·沙赫布OUT医疗城)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22694 2025-10-28 cs.CV cs.CL cs.IR 81%

Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation

Shu Zhao, Tianyi Shen, Nilesh Ahuja, Omesh Tickoo, Vijaykrishnan Narayanan

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Intel(英特尔)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at NeurIPS 2025 UniReps Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20393 2025-10-24 cs.CV cs.MM 81%

Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval

Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim

机构 * Singapore Management University(新加坡管理大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.MM

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20193 2025-10-24 cs.IR cs.CL cs.CV cs.LG 81%

Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures

Rahul Raja, Arpita Vats

机构 * Carnegie Mellon University(卡内基梅隆大学) Boston University(波士顿大学)

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.CL

Comments In Proceedings of the 2nd ACM Workshop in AI-powered Question and Answering Systems (AIQAM '25), October 27-28, 2025, Dublin, Ireland. ACM, New York, NY, USA, 8 pages. https://doi.org/10.1145/3746274.3760393

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14605 2025-10-21 cs.CV cs.AI 81%

Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering

Yuyang Hong, Jiaqi Gu, Qi Yang, Lubin Fan, Yue Wu, Ying Wang, Kun Ding, Shiming Xiang, Jieping Ye

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Alibaba Cloud Computing(阿里巴巴云计算)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14824 2025-10-17 cs.CL cs.CV cs.IR 81%

Supervised Fine-Tuning or Contrastive Learning? Towards Better Multimodal LLM Reranking

Ziqi Dai, Xin Zhang, Mingxin Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Wenjie Li, Min Zhang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03477 2025-09-26 cs.LG cs.AI cs.CV 81%

Robult: Leveraging Redundancy and Modality Specific Features for Robust Multimodal Learning

Duy A. Nguyen, Abhi Kamboj, Minh N. Do

机构 * Siebel School of Computing and Data Science, UIUC, US(UIUC计算机与数据科学学院) Department of Electrical and Computer Engineering, UIUC, US(UIUC电气与计算机工程学院) VinUni-Illinois Smart Health Center, VinUniversity, Vietnam(Vin大学-伊利诺伊智能健康中心)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted and presented at IJCAI 2025 in Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21105 2025-09-23 cs.IR cs.AI cs.CL 81%

AgentMaster: A Multi-Agent Conversational Framework Using A2A and MCP Protocols for Multimodal Information Retrieval and Analysis

Callie C. Liao, Duoduo Liao, Sai Surya Gadiraju

机构 * Stanford University(斯坦福大学) George Mason University(乔治·马歇尔大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15882 2025-09-22 cs.CV cs.AI 81%

Self-Supervised Cross-Modal Learning for Image-to-Point Cloud Registration

Xingmei Wang, Xiaoyu Hu, Chengkai Huang, Ziyan Zeng, Guohao Nie, Quan Z. Sheng, Lina Yao

专题命中 跨模态检索 :cross-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15470 2025-09-22 cs.CV cs.AI 81%

Self-supervised learning of imaging and clinical signatures using a multimodal joint-embedding predictive architecture

Thomas Z. Li, Aravind R. Krishnan, Lianrui Zuo, John M. Still, Kim L. Sandler, Fabien Maldonado, Thomas A. Lasko, Bennett A. Landman

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08338 2025-09-11 cs.CV cs.AI cs.LG 81%

Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis

Jihyun Moon, Charmgil Hong

机构 * Handong Global University(-handong全球大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Medical Image Computing and Computer-Assisted Intervention (MICCAI) ISIC Skin Image Analysis Workshop (MICCAI ISIC) 2025; 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01341 2025-09-03 cs.CV cs.AI 81%

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation

Yunus Serhat Bicakci, Joseph Shingleton, Anahid Basiri

机构 * Vocational School of Social Sciences, Marmara University(马尔马拉大学社会科学职业学校) Geospatial Data Science Group, School of Geographical & Earth Sciences, University of Glasgow(格拉斯哥大学地理与地球科学学院空间数据科学小组)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏