arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 551 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 551 篇

2510.04145 2025-10-07 cs.CV cs.CL cs.IR 84%

Automating construction safety inspections using a multi-modal vision-language RAG framework

Chenxin Wang, Elyas Asadi Shamsabadi, Zhaohui Chen, Luming Shen, Alireza Ahmadian Fard Fini, Daniel Dias-da-Costa

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

Comments 33 pages, 11 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20769 2025-09-26 cs.IR cs.AI cs.CV 84%

Provenance Analysis of Archaeological Artifacts via Multimodal RAG Systems

Tuo Zhang, Yuechun Sun, Ruiliang Liu

机构 * Museus University of Science and Technology of China(中国科学技术大学) British Museum(大英博物馆)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10571 2025-09-23 cs.AI cs.CL 84%

Agentic AI with Orchestrator-Agent Trust: A Modular Visual Classification Framework with Trust-Aware Orchestration and RAG-Based Reasoning

Konstantinos I. Roumeliotis, Ranjan Sapkota, Manoj Karkee, Nikolaos D. Tselikas

机构 * University of the Peloponnese, Department of Informatics and Telecommunications(希腊皮埃蒙特大学信息与电信系) Cornell University, Department of Biological and Environmental Engineering(康奈尔大学生物与环境工程系)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16701 2025-09-03 cs.IR cs.CL 84%

AlzheimerRAG: Multimodal Retrieval Augmented Generation for Clinical Use Cases using PubMed articles

Aritra Kumar Lahiri, Qinmin Vivian Hu

机构 * Department of Computer Science, Toronto Metropolitan University(计算机科学系,多伦多 Metropolitan 大学)

专题命中 多模态RAG :retrieval augmented generation(title);retrieval-augmented generation(abstract);RAG(abstract);分类 cs.IR、cs.CL

Journal ref Machine Learning and Knowledge Extraction. 2025; 7(3):89

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09170 2025-08-14 cs.LG cs.AI cs.CV cs.IR 84%

Multimodal RAG Enhanced Visual Description

Amit Kumar Jaiswal, Haiming Liu, Ingo Frommholz

机构 * Indian Institute of Technology (BHU)(印度理工学院(BHU)) University of Southampton(南安普顿大学) Modul University Vienna(维也纳应用科技大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

Comments Accepted by ACM CIKM 2025. 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16035 2025-07-15 cs.LG cs.AI cs.IR 84%

Vision-Guided Chunking Is All You Need: Enhancing RAG with Multimodal Document Understanding

Vishesh Tripathi, Tanmay Odapally, Indraneel Das, Uday Allu, Biddwan Ahmed

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

Comments 11 pages, 1 Figure, 1 Table

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11063 2025-06-16 cs.CL cs.AI 84%

Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generation

Jiayu Yao, Shenghua Liu, Yiwei Wang, Lingrui Mei, Baolong Bi, Yuyao Ge, Zhecheng Li, Xueqi Cheng

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of California, Merced(加州大学梅德福分校) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07643 2025-04-11 cs.IR cs.CL cs.CV 84%

CollEX -- A Multimodal Agentic RAG System Enabling Interactive Exploration of Scientific Collections

Florian Schneider, Narges Baba Ahmadi, Niloufar Baba Ahmadi, Iris Vogel, Martin Semmann, Chris Biemann

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10886 2025-03-17 cs.CV cs.AI cs.IR cs.LG q-bio.PE 84%

Taxonomic Reasoning for Rare Arthropods: Combining Dense Image Captioning and RAG for Interpretable Classification

Nathaniel Lesperance, Sujeevan Ratnasingham, Graham W. Taylor

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

Comments 12 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00036 2025-02-27 cs.CL cs.AI cs.LG 84%

EMERGE: Enhancing Multimodal Electronic Health Records Predictive Modeling with Retrieval-Augmented Generation

Yinghao Zhu, Changyu Ren, Zixiang Wang, Xiaochen Zheng, Shiyun Xie, Junlan Feng, Xi Zhu, Zhoujun Li, Liantao Ma, Chengwei Pan

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL、cs.AI

Comments CIKM 2024 Full Research Paper; arXiv admin note: text overlap with arXiv:2402.07016

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15040 2025-02-24 cs.CL cs.AI 84%

Reducing Hallucinations of Medical Multimodal Large Language Models with Visual Retrieval-Augmented Generation

Yun-Wei Chu, Kai Zhang, Christopher Malon, Martin Renqiang Min

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL、cs.AI

Comments GenAI4Health - AAAI '25

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05030 2025-01-10 cs.AI cs.CL 84%

A General Retrieval-Augmented Generation Framework for Multimodal Case-Based Reasoning Applications

Ofir Marom

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL、cs.AI

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21943 2024-10-30 cs.CL cs.AI 84%

Beyond Text: Optimizing RAG with Multimodal Inputs for Industrial Applications

Monica Riedler, Stefan Langer

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05131 2024-10-18 cs.LG cs.AI cs.CL cs.CV cs.CY 84%

RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language Models

Peng Xia, Kangyu Zhu, Haoran Li, Hongtu Zhu, Yun Li, Gang Li, Linjun Zhang, Huaxiu Yao

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2024 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12065 2024-07-15 cs.CL cs.AI cs.CY cs.ET cs.MA 84%

RAGAR, Your Falsehood Radar: RAG-Augmented Reasoning for Political Fact-Checking using Multimodal Large Language Models

M. Abdul Khaliq, P. Chang, M. Ma, B. Pflugfelder, F. Miletić

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

Comments 8 pages, submitted to ACL Rolling Review June 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14938 2024-06-24 cs.CL cs.AI 84%

Towards Retrieval Augmented Generation over Large Video Libraries

Yannis Tevissen, Khalil Guetari, Frédéric Petitpont

专题命中 多模态RAG :retrieval augmented generation(title,abstract);RAG(abstract);分类 cs.CL、cs.AI

Comments Accepted in IEEE HSI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17706 2024-05-29 cs.AI cs.CV cs.IR 84%

Video Enriched Retrieval Augmented Generation Using Aligned Video Captions

Kevin Dela Rosa

专题命中 多模态RAG :retrieval augmented generation(title,abstract);RAG(abstract);分类 cs.IR、cs.AI

Comments SIGIR 2024 Workshop on Multimodal Representation and Retrieval (MRR 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22462 2026-02-27 cs.CV cs.IR 84%

MammoWise: Multi-Model Local RAG Pipeline for Mammography Report Generation

MammoWise:多模型本地RAG流水线用于乳腺X线摄影报告生成

Raiyan Jahangir, Nafiz Imtiaz Khan, Amritanand Sudheerkumar, Vladimir Filkov

机构 * University of California, Davis(加州大学戴维斯分校)

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(abstract);分类 cs.IR

AI总结 MammoWise是一种本地多模型流水线,通过多任务分类和检索增强生成技术,实现乳腺X线摄影报告的高准确度生成与分类。

Comments arXiv preprint (submitted 25 Feb 2026). Local multi-model pipeline for mammography report generation + classification using prompting, multimodal RAG (ChromaDB), and QLoRA fine-tuning; evaluates MedGemma, LLaVA-Med, Qwen2.5-VL on VinDr-Mammo and DMID; reports BERTScore/ROUGE-L and classification metrics

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24060 2026-08-13 cs.RO 版本更新 83%

RoboHarness: A Memory-Augmented Policy Harness for Vision-Language-Action Model Robustness via In-Context Adaptation

SOMA:通过上下文适应提升视觉-语言-动作模型鲁棒性的战略编排与内存增强系统

Zhuoran Li, Zhiyang Li, Kaijun Zhou, Jinyu Gu

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract)

AI总结 SOMA通过对比双记忆检索增强生成(RAG)、归因驱动大语言模型(LLM)编排器和可扩展模型上下文协议(MCP)干预,提升视觉-语言-动作模型在分布外任务中的鲁棒性,实验表明其在长周期任务链中提升了89.1%的绝对成功率。

Comments 8 pages, 10 figures, 4 tables. Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026). Project page and source code: https://github.com/LZY-1021/RoboHarness

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16330 2026-07-21 cs.CV 新提交 83%

Local Brushstroke Quality Assessment via Vision-Language Feedback

通过视觉-语言反馈进行局部笔触质量评估

Mio Mitamura, Hirokatsu Kataoka

机构 * Tokyo Institute of Science High School(东京理科大学附属高中) National Institute of Advanced Industrial Science and Technology (AIST)(国立先进工业科学技术研究所) Visual Geometry Group, University of Oxford(牛津大学视觉几何组)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract)

AI总结 研究多模态语言模型能否评估书法局部笔触质量并生成反馈,构建评估框架让三个模型评估图像对并与专家打分比较,还研究了RAG变体,结果显示模型有一定绝对分数准确性,但与专家排名相关性不强,RAG有正负两方面表现。

Comments 6 pages, 8 figures. Accepted to the SAUAFG Workshop at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28128 2026-05-26 cs.LG cs.CR 83%

ORACAL: A Robust and Explainable Multimodal Framework for Smart Contract Vulnerability Detection with Causal Graph Enrichment

ORACAL: 一种基于因果图增强的鲁棒且可解释的智能合约漏洞检测多模态框架

Tran Duong Minh Dai, Triet Huynh Minh Le, M. Ali Babar, Van-Hau Pham, Phan The Duy

机构 * Information Security Lab, University of Information Technology(信息安全部,信息科技大学) Vietnam National University(越南国家大学) School of Computer Science and Information Technology, Adelaide University(计算机科学与信息技术学院,阿德莱德大学)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract)

AI总结 提出ORACAL异构多模态图学习框架,集成控制流图、数据流图和调用图,通过RAG和LLM增强关键子图,并采用因果注意力机制和PGExplainer实现鲁棒且可解释的智能合约漏洞检测。

Comments 21 pages, version 2

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14080 2026-03-03 cs.CY cs.AI 83%

Personalized Education with Generative AI and Digital Twins: VR, RAG, and Zero-Shot Sentiment Analysis for Industry 4.0 Workforce Development

基于生成AI和数字孪生的个性化教育:VR、RAG和零样本情感分析用于工业4.0劳动力发展

Yu-Zheng Lin, Karan Petal, Ahmed H Alhamadah, Sujan Ghimire, Matthew William Redondo, David Rafael Vidal Corona, Jesus Pacheco, Soheil Salehi, Pratik Satam

机构 * Department of Electrical and Computer Engineering, University of Arizona, Tucson, AZ, USA(电气与计算机工程系,亚利桑那大学,图森,亚利桑那州,美国) Department of Systems and Industrial Engineering, University of Arizona, Tucson, AZ, USA(系统与工业工程系,亚利桑那大学,图森,亚利桑那州,美国) Department of Industrial Engineering, University of Sonora, Hermosillo, Mexico(工业工程系,索诺拉大学,赫尔莫索,墨西哥)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本研究提出gAI-PT4I4,利用生成AI和数字孪生技术,结合VR、RAG和零样本情感分析,提升工业4.0劳动力的个性化教育质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00172 2026-03-03 cs.CR cs.AI 83%

Hidden in the Metadata: Stealth Poisoning Attacks on Multimodal Retrieval-Augmented Generation

元数据中的隐藏攻击:多模态检索增强生成中的隐秘污染攻击

Kennedy Edemacu, Mohammad Mahdi Shokri

机构 * The City University of New York, CSI, Staten Island, NY 10314, USA(纽约城市大学,CSI,史泰登岛分校) The City University of New York, Graduate Center, New York, NY 10016, USA(纽约城市大学研究生中心)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.AI

AI总结 本研究提出MM-MEPA,一种通过操纵多模态检索条目元数据来影响模型响应的隐秘污染攻击,展示了其在多模态RAG系统中的高攻击成功率和防御不足问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13179 2026-02-16 cs.IR 83%

Fix Before Search: Benchmarking Agentic Query Visual Pre-processing in Multimodal Retrieval-augmented Generation

在搜索前修复:多模态检索增强生成中代理查询视觉预处理的基准测试

Jiankun Zhang, Shenglai Zeng, Kai Guo, Xinnan Dai, Hui Liu, Jiliang Tang, Yi Chang

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.IR

AI总结 本文提出V-QPP-Bench,通过代理决策任务评估视觉查询预处理的改进,揭示视觉不完美对检索性能的严重影响及训练方法的优化潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08311 2026-01-14 cs.CV cs.AI 83%

Enhancing Image Quality Assessment Ability of LMMs via Retrieval-Augmented Generation

通过检索增强生成提升大型多模态模型的图像质量评估能力

Kang Fu, Huiyu Duan, Zicheng Zhang, Yucheng Zhu, Jun Zhao, Xiongkuo Min, Jia Wang, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Tencent(腾讯)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.AI

AI总结 IQA-RAG通过检索增强生成提升LMMs图像质量评估能力,提供高效替代微调方案

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07329 2026-01-13 cs.CL 83%

BayesRAG: Probabilistic Mutual Evidence Corroboration for Multimodal Retrieval-Augmented Generation

BayesRAG: 基于概率互证的多模态检索增强生成

Xuan Li, Yining Wang, Haocai Luo, Shengping Liu, Jerry Liang, Ying Fu, Weihuang, Jun Yu, Junnan Zhu

机构 * University of Science and Technology of China(中国科学技术大学) Unisound AI Technology Co.Ltd(Unisound AI技术有限公司) MAIS, Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL

AI总结 BayesRAG通过概率互证机制提升多模态检索增强生成的性能,有效解决异构模态的隔离问题。

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03663 2026-01-06 cs.CL cs.CV 83%

UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG

UNIDOC-BENCH: 一个统一的文档中心多模态RAG基准

Xiangyu Peng, Can Qin, Zeyuan Chen, Ran Xu, Caiming Xiong, Chien-Sheng Wu

机构 * Salesforce AI Research(Salesforce人工智能研究)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL

AI总结 UniDoc-Bench是首个大规模文档中心多模态RAG基准,通过多模态问答对评估文本-图像融合与联合检索性能,揭示多模态嵌入不足及视觉上下文补充机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17215 2026-01-06 cs.LG cs.AI cs.CR 83%

How to make Medical AI Systems safer? Simulating Vulnerabilities, and Threats in Multimodal Medical RAG System

如何使医疗AI系统更安全?在多模态医疗RAG系统中模拟漏洞和威胁

Kaiwen Zuo, Zelin Liu, Raman Dutt, Ziyang Wang, Zhongtian Sun, Fan Mo, Pietro Liò

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本文提出MedThreatRAG框架,通过模拟攻击环境揭示医疗RAG系统漏洞,展示跨模态冲突注入对系统性能的严重影响。

Comments Sumbitted to 2026 ICASSP

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19360 2025-12-23 cs.IR 83%

Generative vector search to improve pathology foundation models across multimodal vision-language tasks

生成向量搜索以提升多模态视觉-语言任务中的病理基础模型

Markus Ekvall, Ludvig Bergenstråhle, Patrick Truong, Ben Murrell, Joakim Lundeberg

专题命中 多模态RAG :vector search(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR

AI总结 STHLM通过生成向量搜索方法提升多模态视觉-语言任务中病理基础模型的检索性能,实现10-30%的性能提升和10倍的维度压缩

Comments 13 pages main (54 total), 2 main figures (9 total)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19257 2025-11-25 cs.CR cs.AI cs.LG 83%

Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation

Medusa: 跨模态可转移的对抗攻击用于多模态医疗检索增强生成

Yingjia Shang, Yi Liu, Huimin Wang, Furong Li, Wenfang Sun, Wu Chengyu, Yefeng Zheng

机构 * Westlake University(西湖大学) Heilongjiang University(黑龙江大学) City University of Hong Kong(香港城市大学) Tencent(腾讯)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.AI

AI总结 Medusa提出了一种针对多模态医疗检索增强生成系统的跨模态可转移对抗攻击方法,通过优化扰动和双循环策略实现高攻击成功率并抵御主流防御措施。

Comments Accepted at KDD 2026 First Cycle (full version). Authors marked with * contributed equally. Yi Liu is the lead author

详情

展开后加载摘要…

URL PDF HTML 收藏