Matching Images and Text with Multi-modal Tensor Fusion and Re-ranking
专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV
Comments ACM Multimedia 2019 Oral
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV
Comments ACM Multimedia 2019 Oral
专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);image-text(abstract);分类 cs.CV
Comments 10 pages, 6 figures, Accepted as spotlight at CVPR 2018
先定位再排序:基于知识的视觉问答中无训练实体识别的再探讨
机构 * Rensselaer Polytechnic Institute(伦斯勒理工学院) ; AT&T Chief Data Office(AT&T首席数据办公室)
专题命中 跨模态检索 :MLLM(summary_cn,abstract);multi-modal(abstract);分类 cs.CV、cs.CL
AI总结 针对知识型视觉问答中实体与证据双重定位瓶颈,提出解耦实体识别与段落排序的无训练IBA框架,利用MLLM候选名称选择与文本重排序器,在降低复杂度同时提升性能。
Comments Accepted by ACL 2026 Findings. Project page https://github.com/VAN-QIAN/ACL26-IBA/
基于注意力的视觉文档检索增强
机构 * Alibaba Group(阿里巴巴集团) ; University of Chinese Academy of Sciences(中国科学院大学)
专题命中 跨模态检索 :MLLM(abstract,abstract_cn);multimodal(abstract);multi-modal(abstract);cross-modal(abstract)
AI总结 本文提出AGREE框架,通过多模态大语言模型的注意力机制引导检索器识别相关文档区域,提升细粒度相关性建模,实验表明在ViDoRe V2基准上显著优于传统方法。
Comments Published as a conference paper at SIGIR 2026
glance-or-gaze: 通过强化学习激励大 multimodal 模型适应性聚焦搜索
机构 * Hong Kong University of Science and Technology(香港科技大学)
专题命中 跨模态检索 :multimodal(title_cn,summary_cn);分类 cs.CV、cs.AI
AI总结 本文提出 glanco-or-gaze 框架,通过强化学习使大 multimodal 模型主动规划视觉搜索,通过选择性注视和复杂度自适应强化学习提升复杂视觉查询性能。
Journal ref ACL 2026 Findings
专题命中 跨模态检索 :image-text(title,abstract);cross-modal(abstract,comments);分类 cs.CV、cs.AI
Comments This paper has been accepted to the Main Track of AAAI 2025. It contains 9 pages, 7 figures, and is relevant to the areas of cross-modal retrieval and machine learning. The work presents a novel approach in robust image-text retrieval using a tripartite learning framework
Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, No. 18, pp. 19269-19277, 2025
专题命中 跨模态检索 :multi-modal(title);cross-modal(title);分类 cs.CV、cs.MM
Comments 4 pages, 4 figures
用于特征发现与控制的多模态模型差异分析
机构 * University of Oxford(牛津大学) ; Microsoft(微软公司)
专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 本研究提出MMDiff多模态模型差异分析框架,训练多模态SAEs以识别多模态训练改变的特征,实现特征隔离、检测与控制,在空间、OCR任务及多模态安全攻击评估中展现出良好效果。
Comments Preprint. Accepted at ICML 2026 Trustworthy AI for Good Workshop
UEmbed:统一稀疏与稠密多模态嵌入
机构 * CASIA(中国科学院自动化研究所) ; Alibaba Group(阿里巴巴集团) ; University of Chinese Academy of Sciences(中国科学院大学) ; Yale University(耶鲁大学)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 UEmbed是一种仅解码器的多模态嵌入模型,可单次前向传播生成稀疏与稠密表示,在MMEB-v2和BEIR上表现优异,统一了两类嵌入并支持多模态智能体应用。
Stellar:面向自然语言查询的可扩展多模态文档检索
专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract,abstract_cn)
AI总结 提出Stellar框架,通过磁盘存储令牌级文档嵌入并动态加载候选嵌入,结合词汇表示过滤和高效磁盘支持的后交互,在保持检索效果的同时将内存开销和查询延迟降低1-2个数量级。
多模态跨域对齐网络用于视频时刻检索
机构 * Hubei Key Laboratory of Distributed System Security(湖北分布式系统安全重点实验室) ; Hubei Engineering Research Center on Big Data Security(湖北大数据安全工程研究中心) ; School of Cyber Science and Engineering(网络安全学院) ; Huazhong University of Science and Technology(华中科技大学) ; Wangxuan Institute of Computer Technology(王轩计算机技术研究所) ; Peking University(北京大学) ; School of Computer Science and Technology(计算机科学与技术学院) ; Key Laboratory of Information Storage System Ministry of Education of China(信息存储系统教育部重点实验室)
专题命中 跨模态检索 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
AI总结 提出多模态跨域对齐网络,通过域对齐、跨模态对齐和特定对齐三个模块,解决跨域视频时刻检索中域差异和语义鸿沟问题。
Comments Accepted by IEEE Transactions on Multimedia
乘积交互中的隐藏问题:揭示多模态对比学习中的脆弱性
机构 * Berlin Institute of Health, Charité - Universitätsmedizin Berlin(柏林健康研究所,柏林查理医院) ; Intelligent Medicine Institute, Fudan University(智能医学研究院,复旦大学) ; Department of Mathematics and Computer Science, Freie Universität Berlin(数学与计算机科学系,柏林自由大学)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract)
AI总结 本文提出Gated Symile,通过引入对比门机制,解决多模态对比学习中因单个模态信息不足、错位或缺失导致的脆弱性问题,提升检索准确率。
净化多模态检索:用于RAG的片段级证据选择
专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract,abstract_cn)
AI总结 本文提出FES-RAG框架,通过片段级证据选择提升多模态检索效果,减少噪声干扰,实验显示在M2RAG基准上性能提升27%。
MegaRAG:基于多模态知识图谱的检索增强生成
机构 * Department of Computer Science and Information Engineering, National Taiwan University(国立台湾大学计算机科学与信息工程系) ; Department of Mathematics, National Kaohsiung Normal University(高雄师范大学数学系) ; FinTech Center, National Taiwan University(国立台湾大学金融科技中心) ; E.SUN Financial Holding Co., Ltd.(E.SUN财务公司)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 MegaRAG通过引入多模态知识图谱,提升大语言模型在跨模态推理中的内容理解能力,在文本和多模态语料中均优于现有RAG方法。
Comments ACL 2026
通过多模态大语言模型扩展音频-文本检索
机构 * Visual Geometry Group, University of Oxford(牛津大学视觉几何组) ; Epidemic Sound ; School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)
专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract)
AI总结 AuroLA通过多模态大语言模型实现音频-文本检索的扩展,利用可扩展的数据管道和混合NCE损失提升检索性能。
Comments Technical Report
RAD:迈向可信的检索增强多模态临床诊断
机构 * College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) ; Shanghai AI Laboratory(上海人工智能实验室) ; CMIC, Shanghai Jiao Tong University(上海交通大学计算机学院) ; School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) ; Institute of Artificial Intelligence for Medicine, Shanghai Jiao Tong University(上海交通大学医学人工智能研究所)
专题命中 跨模态检索 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract)
AI总结 RAD通过检索增强多模态模型,提升临床诊断的可信度与准确性,实现任务特定知识的显式注入。
Comments Accepted to NeurIPS 2025
M3DR:迈向通用多语言多模态文档检索
机构 * CognitiveLab(认知实验室)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 M3DR提出了一种通用多语言多模态文档检索框架,通过对比训练实现跨语言和跨模态对齐,显著提升了多语言场景下的检索性能。
机构 * NC AI(NC人工智能)
专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI、cs.MM
Comments Accepted to MMGenSR Workshop (CIKM 2025)
专题命中 跨模态检索 :multi-modal(title,abstract);multimodal(abstract);cross-modal(abstract)
Comments Accepted for publication at EMNLP 2025 Findings. Code and data publicly available at https://github.com/J1mL1/DocMMIR
机构 * Emory University(埃默里大学) ; University of Michigan, Ann Arbor(密歇根大学安娜堡分校)
专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 跨模态检索 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments 28 pages, 6 figures, under review
机构 * National Taiwan University(国立台湾大学)
专题命中 跨模态检索 :cross-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
Comments ICMR 2025
机构 * University of Ottawa(渥太华大学) ; American University(美国美国大学)
专题命中 跨模态检索 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI、cs.MM
Comments Accepted at NAACL 2025 Findings; camera-ready version
Journal ref Findings Assoc. Comput. Linguistics: NAACL 2025, 5201-5215 (2025)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);multimodal foundation model(abstract)
Comments WWW2025 Workshop Summary
专题命中 跨模态检索 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Published as a conference paper at ICLR 2025. The model is available at https://huggingface.co/StevenHH2000/Finedefics
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);image-text(abstract)
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS
专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
Comments EMNLP 2024
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments EMNLP 2024 Industry Track Accepted (Camera-Ready Version)