arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3484 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3484 篇

2202.10401 2022-03-29 cs.CV 70%

Vision-Language Pre-Training with Triple Contrastive Learning

Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi, Junzhou Huang

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments CVPR 2022; code: https://github.com/uta-smile/TCL

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.11499 2022-02-16 cs.SD cs.LG eess.AS 70%

Wav2CLIP: Learning Robust Audio Representations From CLIP

Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, Juan Pablo Bello

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 eess.AS

Comments Copyright 2022 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10807 2021-10-22 cs.CV 70%

Text-Based Person Search with Limited Data

Xiao Han, Sen He, Li Zhang, Tao Xiang

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments 20 pages, 7 figures, 6 tables, to appear in BMVC2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.00883 2021-09-03 cs.CV 70%

MOON: Multi-Hash Codes Joint Learning for Cross-Media Retrieval

Donglin Zhang, Xiao-Jun Wu, He-Feng Yin, Josef Kittler

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.04305 2021-07-07 cs.CV 70%

Learning the Best Pooling Strategy for Visual Semantic Embedding

Jiacheng Chen, Hexiang Hu, Hao Wu, Yuning Jiang, Changhu Wang

专题命中 跨模态检索 :multi-modal(abstract);image-text(abstract);分类 cs.CV

Comments CVPR 2021 camera-ready (oral). The new version fixes a few typos and updates citations

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.07059 2020-01-22 cs.CV cs.CC 70%

Accuracy vs. Complexity: A Trade-off in Visual Question Answering Models

Moshiur R. Farazi, Salman H. Khan, Nick Barnes

专题命中 跨模态检索 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.05930 2019-09-24 cs.CV cs.LG cs.RO 70%

Cross-View Policy Learning for Street Navigation

Ang Li, Huiyi Hu, Piotr Mirowski, Mehrdad Farajtabar

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.03477 2019-08-12 cs.CV 70%

Fine-Grained Action Retrieval Through Multiple Parts-of-Speech Embeddings

Michael Wray, Diane Larlus, Gabriela Csurka, Dima Damen

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted for presentation at ICCV. Project Page: https://mwray.github.io/FGAR

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.06368 2018-08-21 cs.CV 70%

Learning to Learn from Web Data through Deep Semantic Embeddings

Raul Gomez, Lluis Gomez, Jaume Gibert, Dimosthenis Karatzas

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments ECCV MULA Workshop 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.09347 2017-11-28 cs.CV 70%

HashGAN:Attention-aware Deep Adversarial Hashing for Cross Modal Retrieval

Xi Zhang, Siyu Zhou, Jiashi Feng, Hanjiang Lai, Bo Li, Yan Pan, Jian Yin, Shuicheng Yan

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments 10 pages, 8 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.01311 2017-08-07 cs.CV 70%

Automatic Spatially-aware Fashion Concept Discovery

Xintong Han, Zuxuan Wu, Phoenix X. Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, Larry S. Davis

专题命中 跨模态检索 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments ICCV 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05197 2026-03-17 cs.AI cs.CL cs.CV 69%

QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering

QA-Dragon:面向知识密集型视觉问答的查询感知动态RAG系统

Zhuohang Jiang, Pangjing Wu, Xu Yuan, Wenqi Fan, Qing Li

专题命中 跨模态检索 :multimodal(abstract,journal_ref);分类 cs.CV、cs.CL、cs.AI

AI总结 QA-Dragon通过引入领域路由器和搜索路由器,实现多模态、多轮和多跳推理,提升复杂视觉问答任务的推理性能,实验显示其在单源、多源和多轮任务中均优于基线模型。

Comments The source code for our system is released in https://github.com/jzzzzh/QA-Dragon

Journal ref 2025 KDD Cup Workshop for Multimodal Retrieval Augmented Generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24823 2026-08-26 cs.LG q-bio.QM 新提交 67%

BioKERN: Biological Kernel Regularization for Histology-to-Transcriptomics Neighborhood Retrieval

BioKERN:用于组织学-转录组学邻域检索的生物核正则化方法

Seungik Cho, Betul Orcan-Ekmekci

机构 * Rice University(莱斯大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract)

AI总结 该研究针对空间解析生物学的组织学-转录组学邻域检索问题,提出BioKERN框架,通过结合转录组相似性与空间邻近性构建生物核实现正则化,在相关数据集上较BLEEP提升了检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01493 2026-08-24 cs.IR cs.AI cs.CV cs.MM 版本更新 67%

PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval

PhotoBench: 超越视觉匹配,迈向个性化意图驱动的图片检索

Tianyi Xu, Rong Shan, Junjie Wu, Jiadeng Huang, Teng Wang, Jiachen Zhu, Wenteng Chen, Minxin Tu, Quantao Dou, Zhaoxiang Wang, Changwang Zhang, Weinan Zhang, Jun Wang, Jianghao Lin

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 PhotoBench 是首个基于真实个人相册构建的基准,旨在通过多源意图驱动推理提升个性化图片检索能力。

Comments Accepted by KDD'26 Benchmark track

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12627 2026-08-20 cs.CV cs.AI cs.CL cs.HC 版本更新 67%

EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory

EgoCITE:面向长时程自我中心记忆的上下文增强索引与时序感知检索

Le Zhang, Hao Chen, Vlad Roznyatovskiy, Jianzhong Zhang, Ke Sun

机构 * University of Michigan(密歇根大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本研究针对长时程自我中心记忆系统的索引不可靠、忽略时序意图的问题,提出EgoCITE框架,经多数据集评估,其准确率优于基线且成本显著低于长上下文LLM智能体。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01280 2026-08-07 cs.CV cs.AI cs.CL 版本更新 67%

Look Twice: Training-Free Evidence Highlighting for Knowledge-based Visual Question Answering

二次审视:多模态大语言模型中的免训练证据高亮

Marco Morini, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷焦艾米利亚大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出Look Twice框架,通过模型注意力模式识别相关视觉区域和文本证据,提升预训练MLLMs在多模态证据利用中的表现,实验显示在多个VQA基准上效果显著。

Comments Project Page: https://aimagelab.github.io/LoT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00551 2026-08-04 cs.IR 新提交 67%

PHA-Net: Prototype-based Hierarchical Alignment Network for Text-Video Retrieval

PHA-Net:用于文本-视频检索的基于原型的分层对齐网络

Xiaolun Jing, Kezhao Yin, Xinxing Yang, Genke Yang, Jian Chu

专题命中 跨模态检索 :cross-modal(abstract);image-text(abstract)

AI总结 针对文本-视频检索中跨模态语义不匹配及计算成本高的问题,提出PHA-Net,以模态共享原型为桥梁,结合原型支持的令牌合并模块与原型对比损失,在四个基准数据集上取得显著性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23321 2026-07-30 cs.IR 版本更新 67%

MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models

MMEB-V3: 多模态嵌入模型性能差距的测量

Haohang Huang, Xuan Lu, Mingyi Su, Xuan Zhang, Ziyan Jiang, Ping Nie, Kai Zou, Tomas Pfister, Wenhu Chen, Wei Zhang, Xiaoyu Shen, Rui Meng

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract)

AI总结 本文提出MMEB-V3基准,评估文本、图像、视频、音频及基于代理的多模态嵌入,发现现有模型在跨模态检索中存在显著偏差和不足。

Comments Accepted at COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14115 2026-07-17 cs.AI cs.CL cs.CV 新提交 67%

DialogueVPR: Towards Conversational Visual Place Recognition

DialogueVPR:迈向对话式视觉场所识别

Yukun Song, Changwei Wang, Xingtian Pei, Shibiao Xu, Wenhao Xu, Shunpeng Chen, Yu Zhang, Ke Zhang, Rongtao Xu, Xuxiang Feng, Pengyang Wang

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Shandong Computer Science Center, Qilu University of Technology(山东省计算中心(国家超级计算济南中心),齐鲁工业大学) Macquarie University(麦考瑞大学) Spatialtemporal AI(时空人工智能公司) University of Macau(澳门大学) Aerospace Information Research Institute(航天信息研究所)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 研究针对语言引导地理定位中现有方法不足,提出DlgPR,将场所识别转为对话驱动推理过程。构建DlgQuest-Cities基准及统一推理框架,用课程训练DQ-pilot,通过特定指标指导学习并实验,该方法显著优于基线。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30196 2026-06-30 cs.CL cs.AI cs.LG eess.AS 67%

Forewarned is Forearmed: When Non-Sequential Embedding Turns Into an Anomaly Detector

有备无患:当非序列嵌入变成异常检测器

Elys Allesiardo, Antoine Caubrière, Valentin Vielzeuf

机构 * Orange Research(Orange研究院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 本文深入分析非序列多模态句子级嵌入(SONAR模型),发现某些嵌入维度对扰动敏感,可作为解码异常指标,并利用编解码一致性构建准确检测器,同时探索修正异常维度。

Comments Accepted for presentation at LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29069 2026-06-30 cs.AI cs.CL cs.CV 67%

Low-cost concept-based localized explanations: How far can we get with training-free approaches?

低成本基于概念的可解释性:无训练方法能走多远?

Darian Fernández-Gutiérrez, Rafael Bello, Marilyn Bello, Natalia Díaz-Rodríguez

机构 * Dept. of Computer Science and Artificial Intelligence, University of Granada (UGR)(计算机科学与人工智能系,格拉纳达大学(UGR))

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出零样本概念命名协议,利用中等规模多模态大模型对局部区域进行概念标注,无需训练即可实现62%-88%的物体级精确匹配,为低成本可解释AI提供新思路。

Comments 6 pages, 2 figures, 4 tables. Accepted at the 2026 IEEE International Conference on Artificial Intelligence (CAI), 8-10 May 2026, Granada, Spain. Code: https://github.com/darianfgUgr/CoNa

Journal ref 2026 IEEE International Conference on Artificial Intelligence (CAI), Granada, Spain, 2026, pp. 1405-1410

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28344 2026-06-30 cs.IR cs.AI cs.CL cs.CV cs.LG 67%

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation

PIXELRAG:网页截图在检索增强生成中优于文本

Yichuan Wang, Zhifei Li, Zirui Wang, Paul Teiletche, Lesheng Jin, Matei Zaharia, Joseph E. Gonzalez, Sewon Min

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出PixelRAG方法,以网页截图替代文本进行检索和阅读,利用视觉嵌入模型和对比学习,在30M截图库上实现端到端RAG,在多项任务中优于文本基线,准确率提升最高18.1%,并通过图像压缩降低3倍token成本。

Comments Our code is available at https://github.com/StarTrail-org/PixelRAG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10020 2026-06-05 cs.CL cs.AI cs.CV 67%

The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?

性能提升的幻象:为何对比解码无法减轻多模态大语言模型中的对象幻觉?

Hao Yin, Guangzong Si, Zilei Wang

机构 * University of Science and Technology of China(中国科学技术大学) Eastern Institute of Technology, Ningbo(宁波东部技术研究所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文研究了对比解码方法在减轻多模态大语言模型(MLLMs)中对象幻觉方面的有效性,发现其性能提升主要源于两个误导性因素,挑战了对比解码策略的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27009 2026-05-27 cs.LG 67%

SCENT: Aligning Mass Spectra with Molecular Structure for Olfactory Perception

SCENT: 将质谱与分子结构对齐用于嗅觉感知

Ziqi Zhang, Eunyeong Jin, Miguel Vasco, Farzaneh Taleb, Nona Rajabi, Alexandra Gutmann, Jonathan Williams, Antônio H. Ribeiro, Danica Kragic

机构 * Dept. of Intelligent Systems, KTH Royal Institute of Technology(智能系统系,皇家理工学院) Atmospheric Chemistry Dept., Max Planck Institute for Chemistry(大气化学部,马克斯·普朗克研究所) Dept. of Information Technology, Uppsala University(信息科技系,乌普萨拉大学) Science for Life Laboratory (SciLifeLab), Uppsala(生命科学实验室(SciLifeLab),乌普萨拉)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract)

AI总结 提出SCENT多模态对比学习框架,通过将电子电离质谱表示与预训练化学结构嵌入对齐,在无需分子结构的情况下实现与结构模型相当的嗅觉预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22715 2026-05-26 cs.CV cs.AI cs.CL cs.HC 67%

AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild

AnyMo:野外人体运动的几何感知与设置无关建模

Baiyu Chen, Zechen Li, Wilson Wongso, Lihuan Li, Xiachong Lin, Hao Xue, Benjamin Tag, Flora Salim

机构 * The University of New South Wales(新南威尔士大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出AnyMo框架,通过物理模拟生成多样化IMU信号、图编码器预训练和LLM对齐,实现跨设备/数据集的零样本活动识别、跨模态检索和运动描述,性能显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26467 2026-04-30 cs.CR 67%

Differentially Private Contrastive Learning via Bounding Group-level Contribution

差分隐私对比学习通过限制群体级贡献

Kecen Li, Chen Gong, Zinan Lin, Tianhao Wang, Xiaokui Xiao

专题命中 跨模态检索 :multi-modal(abstract);image-text(abstract)

AI总结 本文提出DP-GCL框架,通过限制梯度依赖提升差分隐私对比学习效果,实验显示在多个数据集上均取得最佳性能,图像分类准确率提升5.6%,图像-文本检索准确率提升20.1%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18123 2026-04-03 cs.CV cs.AI cs.CL cs.LG 67%

Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models

偏差是子空间,而非坐标:视觉-语言模型中事后去偏的几何重思

Dachuan Zhao, Weiyue Li, Zhenda Shen, Yushu Qiu, Bowen Xu, Haoyu Chen, Yongchao Chen

机构 * Harvard University(哈佛大学) MIT(麻省理工学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出SPD框架,通过几何方法识别并去除线性可解码的偏子空间,提升视觉-语言模型的公平性与任务性能。

Comments Accepted at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16737 2026-03-18 cs.CV cs.AI cs.CL 67%

Retrieving Counterfactuals Improves Visual In-Context Learning

检索反事实改进视觉上下文学习

Guangzhi Xiong, Sanchit Sinha, Zhenghao He, Aidong Zhang

机构 * University of Virginia(弗吉尼亚大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出CIRCLES框架,通过主动检索反事实样例提升视觉语言模型的因果推理能力,实验表明其在多个数据集上优于现有方法,尤其在小规模模型和信息稀缺场景下表现突出。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02556 2026-03-04 cs.CV cs.AI cs.CL cs.LG 67%

Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs

通过对比的视角:VLMs中的自改进视觉推理

Zhiyu Pan, Yizheng Wu, Jiashen Hua, Junyi Feng, Shaotian Yan, Bing Deng, Zhiguo Cao, Jieping Ye

机构 * Huazhong University of Science and Technology(华中科技大学) Alibaba Cloud(阿里云)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 通过视觉对比提升VLMs的推理能力,提出VC-STaR框架,有效减少推理幻觉并提升多种VLMs的视觉推理性能。

Comments 19 pages, 9 figures, accepted to ICLR 2026 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01990 2026-03-03 cs.AI cs.CL cs.CV 67%

According to Me: Long-Term Personalized Referential Memory QA

根据我:长期个性化参照记忆问答

Jingbiao Mei, Jinghong Chen, Guangyu Yang, Xinyu Hou, Margaret Li, Bill Byrne

机构 * Department of Engineering, University of Cambridge, United Kingdom(剑桥大学工程系) Department of Physics, University of Cambridge, United Kingdom(剑桥大学物理系) Independent Researcher(独立研究者)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本文提出ATM-Bench,首个多模态多来源个性化参照记忆问答基准,并提出Schema-Guided Memory方法提升记忆推理性能

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏