M3Retrieve: Benchmarking Multimodal Retrieval for Medicine
机构 * Indian Institute of Technology Patna(印度理工学院帕纳瓦分校) ; King Mongkut’s Institute of Technology Ladkrabang(拉差丹awan技术大学)
专题命中 多模态RAG :RAG(abstract);分类 cs.IR、cs.AI
Comments EMNLP Mains 2025
AI 大模型
检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。
机构 * Indian Institute of Technology Patna(印度理工学院帕纳瓦分校) ; King Mongkut’s Institute of Technology Ladkrabang(拉差丹awan技术大学)
专题命中 多模态RAG :RAG(abstract);分类 cs.IR、cs.AI
Comments EMNLP Mains 2025
机构 * Department of Education and Humanities, University of Modena and Reggio Emilia(教育与人文学院, Modena and Reggio Emilia大学)
专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI
机构 * Amazon(亚马逊)
专题命中 多模态RAG :RAG(abstract);分类 cs.IR、cs.AI
Comments Conference: KDD conference workshop: https://kdd-eval-workshop.github.io/genai-evaluation-kdd2025/
专题命中 多模态RAG :RAG(abstract);分类 cs.CL、cs.AI
Comments V4, fixes to title and formatting
机构 * Leiden University(莱顿大学)
专题命中 多模态RAG :dense retrieval(abstract);分类 cs.IR、cs.AI
Comments Accepted at ACL 2025, Main track. 13 Pages, 1 figure
机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中 多模态RAG :retriever(abstract);分类 cs.IR、cs.CL
Comments 18 pages. Code and data: https://github.com/meetdavidwan/clamr
专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI
专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI
Comments ICLR 2025
专题命中 多模态RAG :retriever(abstract);分类 cs.CL、cs.AI
Comments CVPR 2025
专题命中 多模态RAG :retrieval augmented generation(abstract);分类 cs.CL、cs.AI
Comments Published in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (EMNLP 2024) Best Demo Paper Award at EMNLP 2024
Journal ref EMNLP 2024 (System Demonstrations), pp. 46-52
专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.CL、cs.AI
专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.IR、cs.AI
Comments Code is available at https://github.com/cnzzx/VSA
专题命中 多模态RAG :retrieval augmented generation(abstract);分类 cs.CL、cs.AI
专题命中 多模态RAG :retriever(abstract);分类 cs.AI、cs.DB
Comments Governance, Understanding and Integration of Data for Effective and Responsible AI (GUIDE-AI '24), June 14, 2024, Santiago, AA, Chile
专题命中 多模态RAG :vector search(abstract);分类 cs.IR、cs.DB
Comments This paper has been accepted by ICDE 2024
专题命中 多模态RAG :retriever(abstract);分类 cs.CL、cs.AI
Comments CVPR 2023
专题命中 多模态RAG :RAG(abstract);分类 cs.CL、cs.AI
Comments Accepted to EMNLP 2022 main conference
专题命中 多模态RAG :retriever(abstract);分类 cs.CL、cs.AI
Comments Accepted by ACL 2021 main conference
是看不见还是不知道?归因视觉语言模型中的错误
机构 * MBZUAI The University of Melbourne(MBZUAI墨尔本大学)
专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.CL
AI总结 研究视觉语言模型在回答需额外知识问题时的错误,提出统一框架分离失败模式,探讨预生成信号能否预测错误源,发现可在解码前预测,能据此进行针对性干预。
通过检索增强的可靠性感知推理缓解多模态系统中的视觉幻觉
机构 * University of Massachusetts, Dartmouth(马萨诸塞大学达特茅斯分校)
专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI
AI总结 提出一种检索增强的可靠性感知推理框架,利用外部视觉证据库和多个可靠性指标进行决策门控,在不重训练模型的情况下减少视觉幻觉,将接受预测准确率从85.84%提升至88.88%。
Comments 29 pages, 9 figures
显著知识路径:用于高效知识密集型多模态问答的稀疏跨模态路由
专题命中 多模态RAG :dense retrieval(abstract);分类 cs.AI
AI总结 研究知识密集型多模态问答,提出SKIP架构,通过问题引导视觉令牌修剪等方法,沿稀疏路径计算路由,结合自适应预算控制器,在五个基准测试中,以更少计算量和更低延迟达到或超越密集基线准确性。
Comments Accepted at the 43rd International Conference on Machine Learning (ICML 2026) Workshop on Efficient Multimodal Question Answering (EMM-QA), Seoul, South Korea. Copyright 2026 by the author(s). (Archival)
ViMax: 智能体视频生成
机构 * The University of Hong Kong(香港大学) ; South China University of Technology(华南理工大学) ; Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI
AI总结 提出ViMax框架,通过多智能体协作实现长视频生成,利用分层叙事引擎和视觉一致性机制,保证叙事连贯性和视觉一致性。
Comments 20 pages, 13 figures
MM-IssueLoc:用于评估多模态仓库级问题定位中视觉证据的可控基准
机构 * Tsinghua University(清华大学)
专题命中 多模态RAG :retriever(abstract);分类 cs.AI
AI总结 研究针对仓库级问题定位多为文本任务,视觉证据作用不明的情况,引入MM-IssueLoc基准和评估协议,含多语言实例与标注等。评估LLM和检索系统,发现现有系统距可靠多模态定位有差距,该基准让视觉证据成评估变量,助于后续研究。
ManimAgent: 用于视觉教育的自进化多模态智能体
机构 * University of Alberta(阿尔伯塔大学) ; Southeast University(东南大学) ; Virginia Tech(弗吉尼亚理工学院) ; Xidian University(西安电子科技大学) ; Vivavia Inc(Vivavia公司)
专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI
AI总结 提出ManimAgent,通过双通道情节记忆库跨任务传递反思经验,无需权重更新或人工种子,在代码生成任务中提升通过率并减少反思轮次。
Comments Project page: https://manimagent.github.io/. Code: https://github.com/jwj1342/Paper2Manim
阅读,而非思考:理解并弥合多模态大语言模型中文本变为像素时的模态差距
机构 * Johns Hopkins University(约翰霍普金斯大学) ; Amazon(亚马逊) ; New York University(纽约大学) ; Texas A&M University(德克萨斯大学)
专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.CL
AI总结 本文系统诊断多模态大语言模型在处理图像文本时的模态差距,发现其源于模型推理意愿不足而非感知失败,并提出一种轻量级自蒸馏方法有效弥合该差距。
WISE:一种用于视觉场景、音频、物体、人脸、语音和元数据的多模态搜索引擎
机构 * Engineering Science University of Oxford(工程科学大学牛津)
专题命中 多模态RAG :vector search(abstract);分类 cs.IR
AI总结 提出WISE开源多模态搜索引擎,整合场景级和物体级的自然语言与反向图像查询、人脸搜索、音频事件检索、语音转录搜索及元数据过滤,支持跨模态组合查询,采用向量搜索实现高效扩展,可本地部署。
Comments Software: https://www.robots.ox.ac.uk/~vgg/software/wise/ , Online demos: https://www.robots.ox.ac.uk/~vgg/software/wise/demo/ , Example Queries: https://www.robots.ox.ac.uk/~vgg/software/wise/examples/
Journal ref International ACM SIGIR Conference on Research and Development in Information Retrieval (2026)
基于知识的LLM决策支持系统:用于激光粉末床融合的可解释缺陷分析与缓解指导
机构 * Department of Mechanical Engineering, University of Massachusetts Dartmouth(达特茅斯大学机械工程系)
专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.AI
AI总结 本文提出一种整合结构化缺陷知识与LLM推理的知识驱动决策支持系统,用于制造业中激光粉末床融合的可解释缺陷诊断与缓解指导。系统基于包含27种已知缺陷类型的知识库,支持模糊自然语言查询、文献支持的缺陷解释及基于编码工艺知识的缺陷原因和缓解策略指导。
Comments 28 pages, 15 figures
AnalogRetriever: 为模拟电路检索学习跨模态表示
机构 * Tsinghua University(清华大学) ; The University of Hong Kong(香港大学) ; University of Cambridge(剑桥大学) ; Nanjing University of Posts and Telecommunications(南京邮电大学)
专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI
AI总结 本文提出AnalogRetriever,通过构建高质量数据集和三模态检索框架,实现跨模态的模拟电路检索,实验表明其在六个方向上的Recall@1达到75.2%,显著优于现有方法。
Comments 10 pages, 7 figures. Yihan Wang and Lei Li contributed equally to this paper
Re:Verse -- 能读懂漫画吗?
机构 * University of Central Florida(中央佛罗里达大学) ; Indian Institute of Technology, Jodhpur(印度理工学院,朱达浦尔) ; Indian Institute of Technology, Varanasi(印度理工学院,瓦拉纳西)
专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.CL
AI总结 本文通过分析漫画叙事理解,揭示现有VLM在时间因果和跨面板连贯性上的不足,提出新的评估框架,系统研究长篇叙事理解能力。
Comments Accepted (oral) at ICCV (AISTORY Workshop) 2025
Journal ref 2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), pp. 3820-3830
基于MLLM的视觉丰富文档理解综述:方法、挑战与新兴趋势
机构 * The University of Western Australia(西澳大学) ; The University of Melbourne(墨尔本大学) ; Weill Cornell Medicine(韦尔·柯尔医学中心)
专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI
AI总结 本文综述了基于MLLM的视觉丰富文档理解最新进展,探讨了文本、视觉和布局特征的表示与整合技术,以及预训练、指令微调等训练方法,分析了数据稀缺、多页文档处理等挑战及新兴趋势。
Comments Accepted at ACL 2026 Findings