arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 8587 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 552 篇

2603.28128 2026-05-26 cs.LG cs.CR 83%

ORACAL: A Robust and Explainable Multimodal Framework for Smart Contract Vulnerability Detection with Causal Graph Enrichment

ORACAL: 一种基于因果图增强的鲁棒且可解释的智能合约漏洞检测多模态框架

Tran Duong Minh Dai, Triet Huynh Minh Le, M. Ali Babar, Van-Hau Pham, Phan The Duy

机构 * Information Security Lab, University of Information Technology(信息安全部,信息科技大学) Vietnam National University(越南国家大学) School of Computer Science and Information Technology, Adelaide University(计算机科学与信息技术学院,阿德莱德大学)

专题命中 多模态RAG :RAG(summary_cn,abstract);retrieval-augmented generation(abstract)

AI总结 提出ORACAL异构多模态图学习框架,集成控制流图、数据流图和调用图,通过RAG和LLM增强关键子图,并采用因果注意力机制和PGExplainer实现鲁棒且可解释的智能合约漏洞检测。

Comments 21 pages, version 2

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14080 2026-03-03 cs.CY cs.AI 83%

Personalized Education with Generative AI and Digital Twins: VR, RAG, and Zero-Shot Sentiment Analysis for Industry 4.0 Workforce Development

基于生成AI和数字孪生的个性化教育:VR、RAG和零样本情感分析用于工业4.0劳动力发展

Yu-Zheng Lin, Karan Petal, Ahmed H Alhamadah, Sujan Ghimire, Matthew William Redondo, David Rafael Vidal Corona, Jesus Pacheco, Soheil Salehi, Pratik Satam

机构 * Department of Electrical and Computer Engineering, University of Arizona, Tucson, AZ, USA(电气与计算机工程系,亚利桑那大学,图森,亚利桑那州,美国) Department of Systems and Industrial Engineering, University of Arizona, Tucson, AZ, USA(系统与工业工程系,亚利桑那大学,图森,亚利桑那州,美国) Department of Industrial Engineering, University of Sonora, Hermosillo, Mexico(工业工程系,索诺拉大学,赫尔莫索,墨西哥)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本研究提出gAI-PT4I4,利用生成AI和数字孪生技术,结合VR、RAG和零样本情感分析,提升工业4.0劳动力的个性化教育质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00172 2026-03-03 cs.CR cs.AI 83%

Hidden in the Metadata: Stealth Poisoning Attacks on Multimodal Retrieval-Augmented Generation

元数据中的隐藏攻击:多模态检索增强生成中的隐秘污染攻击

Kennedy Edemacu, Mohammad Mahdi Shokri

机构 * The City University of New York, CSI, Staten Island, NY 10314, USA(纽约城市大学,CSI,史泰登岛分校) The City University of New York, Graduate Center, New York, NY 10016, USA(纽约城市大学研究生中心)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.AI

AI总结 本研究提出MM-MEPA,一种通过操纵多模态检索条目元数据来影响模型响应的隐秘污染攻击,展示了其在多模态RAG系统中的高攻击成功率和防御不足问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13179 2026-02-16 cs.IR 83%

Fix Before Search: Benchmarking Agentic Query Visual Pre-processing in Multimodal Retrieval-augmented Generation

在搜索前修复:多模态检索增强生成中代理查询视觉预处理的基准测试

Jiankun Zhang, Shenglai Zeng, Kai Guo, Xinnan Dai, Hui Liu, Jiliang Tang, Yi Chang

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.IR

AI总结 本文提出V-QPP-Bench,通过代理决策任务评估视觉查询预处理的改进,揭示视觉不完美对检索性能的严重影响及训练方法的优化潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08311 2026-01-14 cs.CV cs.AI 83%

Enhancing Image Quality Assessment Ability of LMMs via Retrieval-Augmented Generation

通过检索增强生成提升大型多模态模型的图像质量评估能力

Kang Fu, Huiyu Duan, Zicheng Zhang, Yucheng Zhu, Jun Zhao, Xiongkuo Min, Jia Wang, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Tencent(腾讯)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.AI

AI总结 IQA-RAG通过检索增强生成提升LMMs图像质量评估能力,提供高效替代微调方案

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07329 2026-01-13 cs.CL 83%

BayesRAG: Probabilistic Mutual Evidence Corroboration for Multimodal Retrieval-Augmented Generation

BayesRAG: 基于概率互证的多模态检索增强生成

Xuan Li, Yining Wang, Haocai Luo, Shengping Liu, Jerry Liang, Ying Fu, Weihuang, Jun Yu, Junnan Zhu

机构 * University of Science and Technology of China(中国科学技术大学) Unisound AI Technology Co.Ltd(Unisound AI技术有限公司) MAIS, Institute of Automation, Chinese Academy of Sciences(自动化研究所,中国科学院)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL

AI总结 BayesRAG通过概率互证机制提升多模态检索增强生成的性能,有效解决异构模态的隔离问题。

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03663 2026-01-06 cs.CL cs.CV 83%

UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG

UNIDOC-BENCH: 一个统一的文档中心多模态RAG基准

Xiangyu Peng, Can Qin, Zeyuan Chen, Ran Xu, Caiming Xiong, Chien-Sheng Wu

机构 * Salesforce AI Research(Salesforce人工智能研究)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL

AI总结 UniDoc-Bench是首个大规模文档中心多模态RAG基准,通过多模态问答对评估文本-图像融合与联合检索性能,揭示多模态嵌入不足及视觉上下文补充机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17215 2026-01-06 cs.LG cs.AI cs.CR 83%

How to make Medical AI Systems safer? Simulating Vulnerabilities, and Threats in Multimodal Medical RAG System

如何使医疗AI系统更安全?在多模态医疗RAG系统中模拟漏洞和威胁

Kaiwen Zuo, Zelin Liu, Raman Dutt, Ziyang Wang, Zhongtian Sun, Fan Mo, Pietro Liò

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本文提出MedThreatRAG框架,通过模拟攻击环境揭示医疗RAG系统漏洞,展示跨模态冲突注入对系统性能的严重影响。

Comments Sumbitted to 2026 ICASSP

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19360 2025-12-23 cs.IR 83%

Generative vector search to improve pathology foundation models across multimodal vision-language tasks

生成向量搜索以提升多模态视觉-语言任务中的病理基础模型

Markus Ekvall, Ludvig Bergenstråhle, Patrick Truong, Ben Murrell, Joakim Lundeberg

专题命中 多模态RAG :vector search(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR

AI总结 STHLM通过生成向量搜索方法提升多模态视觉-语言任务中病理基础模型的检索性能,实现10-30%的性能提升和10倍的维度压缩

Comments 13 pages main (54 total), 2 main figures (9 total)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19257 2025-11-25 cs.CR cs.AI cs.LG 83%

Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation

Medusa: 跨模态可转移的对抗攻击用于多模态医疗检索增强生成

Yingjia Shang, Yi Liu, Huimin Wang, Furong Li, Wenfang Sun, Wu Chengyu, Yefeng Zheng

机构 * Westlake University(西湖大学) Heilongjiang University(黑龙江大学) City University of Hong Kong(香港城市大学) Tencent(腾讯)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.AI

AI总结 Medusa提出了一种针对多模态医疗检索增强生成系统的跨模态可转移对抗攻击方法,通过优化扰动和双循环策略实现高攻击成功率并抵御主流防御措施。

Comments Accepted at KDD 2026 First Cycle (full version). Authors marked with * contributed equally. Yi Liu is the lead author

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16654 2025-11-25 cs.CL 83%

Comparison of Text-Based and Image-Based Retrieval in Multimodal Retrieval Augmented Generation Large Language Model Systems

多模态检索增强生成大语言模型系统中基于文本和基于图像的检索比较

Elias Lumer, Alex Cardenas, Matt Melich, Myles Mason, Sara Dieter, Vamse Kumar Subbiah, Pradeep Honaganahalli Basavaraju, Roberto Hernandez

机构 * PricewaterhouseCoopers U.S.(普华永道美国公司)

专题命中 多模态RAG :retrieval augmented generation(title);retrieval-augmented generation(abstract);RAG(abstract);分类 cs.CL

AI总结 本文比较了多模态RAG系统中基于文本和基于图像的检索方法,发现直接多模态嵌入检索在性能和准确性上优于基于LLM总结的方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13093 2025-11-25 cs.CV cs.AI 83%

Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension

Video-RAG: 基于视觉对齐的检索增强长视频理解

Yongdong Luo, Xiawu Zheng, Guilin Li, Shukang Yin, Haojia Lin, Chaoyou Fu, Jinfa Huang, Jiayi Ji, Fei Chao, Jiebo Luo, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学) Nanjing University(南京大学) University of Rochester(罗切斯特大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 Video-RAG通过视觉对齐的辅助文本提升长视频理解性能,无需训练且兼容性强,显著优于专有模型。

Comments Accepted at NeurIPS 2025. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04536 2025-10-07 cs.GR cs.AI cs.CV 83%

3Dify: a Framework for Procedural 3D-CG Generation Assisted by LLMs Using MCP and RAG

Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Satoshi Ohshima, Takahiro Katagiri

机构 * Graduate School of Informatics(信息科学研究生学校) Nagoya University(名古屋大学) Research Institute for Information Technology(信息科学技术研究所)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21556 2025-09-29 cs.CL 83%

VAT-KG: Knowledge-Intensive Multimodal Knowledge Graph Dataset for Retrieval-Augmented Generation

Hyeongcheol Park, Jiyoung Seo, MinHyuk Jang, Hogun Park, Ha Dam Baek, Gyusam Chang, Hyeonsoo Im, Sangpil Kim

机构 * Korea University(韩国大学) Sungkyunkwan University(全北大学) Hanwha Systems(韩华系统)

专题命中 多模态RAG :retrieval-augmented generation(title);retrieval augmented generation(abstract);RAG(abstract);分类 cs.CL

Comments Project Page: https://vatkg.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07600 2025-06-10 cs.CV cs.AI 83%

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding

Nianbo Zeng, Haowen Hou, Fei Richard Yu, Si Shi, Ying Tiffany He

机构 * Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室) College of Computer Science and Software Engineering(计算机科学与软件工程学院)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07399 2025-06-10 cs.CV cs.AI 83%

MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems

Peiru Yang, Jinhua Yin, Haoran Zheng, Xueying Bai, Huili Wang, Yufei Sun, Xintian Li, Shangguang Wang, Yongfeng Huang, Tao Qi

机构 * Tsinghua University(清华大学) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01222 2025-05-23 cs.CV cs.CL 83%

Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG

Wenbin Wang, Yongcheng Jing, Liang Ding, Yingjie Wang, Li Shen, Yong Luo, Bo Du, Dacheng Tao

机构 * Wuhan University(武汉大学) Nanyang Technological University(南洋理工大学) The University of Sydney(悉尼大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13828 2025-05-21 cs.AI 83%

Multimodal RAG-driven Anomaly Detection and Classification in Laser Powder Bed Fusion using Large Language Models

Kiarash Naghavi Khanghah, Zhiling Chen, Lela Romeo, Qian Yang, Rajiv Malhotra, Farhad Imani, Hongyi Xu

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

Comments ASME 2025 International Design Engineering Technical Conferences and Computers and Information in Engineering Conference IDETC/CIE2025, August 17-20, 2025, Anaheim, CA (IDETC2025-168615)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12663 2025-03-18 cs.CV cs.CL cs.LG cs.RO 83%

Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding

Imran Kabir, Md Alimoor Reza, Syed Billah

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13085 2025-03-04 cs.LG cs.CL cs.CV 83%

MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models

Peng Xia, Kangyu Zhu, Haoran Li, Tianze Wang, Weijia Shi, Sheng Wang, Linjun Zhang, James Zou, Huaxiu Yao

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.CL

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00239 2024-12-03 cs.SE cs.AI 83%

Generating a Low-code Complete Workflow via Task Decomposition and RAG

Orlando Marquez Ayala, Patrice Béchard

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

Comments Under review; 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11321 2024-10-16 cs.CL 83%

Self-adaptive Multimodal Retrieval-Augmented Generation

Wenjia Zhai

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);RAG(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12309 2024-08-20 cs.CV cs.IR cs.LG 83%

iRAG: Advancing RAG for Videos with an Incremental Approach

Md Adnan Arefeen, Biplob Debnath, Md Yusuf Sarwar Uddin, Srimat Chakradhar

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.IR

Comments Accepted in CIKM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07016 2024-02-13 cs.AI 83%

REALM: RAG-Driven Enhancement of Multimodal Electronic Health Records Analysis via Large Language Models

Yinghao Zhu, Changyu Ren, Shiyun Xie, Shukai Liu, Hangyuan Ji, Zixiang Wang, Tao Sun, Long He, Zhoujun Li, Xi Zhu, Chengwei Pan

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28002 2026-06-29 cs.CL cs.AI eess.AS 新提交 82%

Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection

对话到检测:用于保险欺诈检测的多模态混合 NLP 流水线

Muhammad Shakeel Akram, Amal Htait, Abdul Hamid Sadka, Emma Meisingseth, Karishma Jaitly

机构 * Aston University(阿斯顿大学) Domestic & General

专题命中 多模态RAG :RAG(summary_cn,abstract);分类 cs.CL、cs.AI

AI总结 提出一种合成多模态框架,结合ASR、NER、LLM-RAG和说话人嵌入,通过规则风险评分检测保险欺诈中的叙述复用、结构不一致和跨案件语音重复。

Comments 10 pages, 8 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05818 2026-04-15 cs.CV cs.CL cs.IR 82%

WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering

WikiSeeker: 重新思考视觉语言模型在基于知识的视觉问答中的作用

Yingjian Zhu, Xinming Wang, Kun Ding, Ying Wang, Bin Fan, Shiming Xiang

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室 (MAIS))

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);retriever(abstract);分类 cs.IR、cs.CL

AI总结 本文提出WikiSeeker框架,通过引入多模态检索器和重新定义视觉语言模型的角色,提升多模态检索性能和答案质量,实现在EVQA、InfoSeek和M2KR数据集上的最优表现。

Comments Accepted by ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03967 2026-03-05 cs.CV 82%

UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted Optimization

UniRain: 基于RAG的数据蒸馏和多目标重加权优化的统一图像去雨

Qianfeng Yang, Qiyuan Guan, Xiang Chen, Jiyu Jin, Guiyue Jin, Jiangxin Dong

机构 * Dalian Polytechnic University(大连理工大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(abstract)

AI总结 UniRain通过基于RAG的数据蒸馏和多目标重加权优化,实现统一图像去雨,有效应对不同雨损场景。

Comments Accepted by CVPR 2026; Project Page: https://github.com/QianfengY/UniRain

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00511 2026-03-03 cs.CV cs.LG 82%

Multimodal Adaptive Retrieval Augmented Generation through Internal Representation Learning

多模态自适应检索增强生成通过内部表示学习

Ruoshuang Du, Xin Sun, Qiang Liu, Bowen Song, Zhongqi Chen, Weiqiang Wang, Liang Wang

机构 * Shanghaitech University School of Information Science(上海科技大学信息科学学院) Chinese Academy of Sciences Institute of Automation(中国科学院自动化研究所)

专题命中 多模态RAG :retrieval augmented generation(title,abstract);RAG(abstract)

AI总结 本文提出MMA-RAG,通过动态评估模型内部知识置信度,提升视觉问答系统在多模态场景下的检索增强生成性能。

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15650 2026-02-18 cs.CV 82%

Concept-Enhanced Multimodal RAG: Towards Interpretable and Accurate Radiology Report Generation

概念增强的多模态RAG:迈向可解释且准确的放射科报告生成

Marco Salmè, Federico Siciliano, Fabrizio Silvestri, Paolo Soda, Rosa Sicilia, Valerio Guarrasi

机构 * Department of Engineering(工程系) Research Unit of Artificial Intelligence and Computer Systems(人工智能与计算机系统研究单位) Università Campus Bio-Medico of Roma(罗马大学生物医学校园) Department of Computer, Control and Management Engineering(计算机、控制与管理工程系) Sapienza University of Rome(罗马萨皮恩扎大学) Department of Diagnostics and Intervention, Radiation Physics, Biomedical Engineering(诊断与介入、辐射物理、生物医学工程系) Umeå University(乌梅拉大学) UniCamillus-Saint Camillus International University of Health Sciences(UniCamillus-圣卡米卢斯国际健康科学大学)

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract)

AI总结 概念增强的多模态RAG通过分解视觉表示为可解释的临床概念,提升放射科报告生成的可解释性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00030 2026-02-10 cs.LG 82%

RAPTOR-AI for Disaster OODA Loop: Hierarchical Multimodal RAG with Experience-Driven Agentic Decision-Making

RAPTOR-AI用于灾难OODA循环:基于经验驱动的多模态RAG框架

Takato Yasuno

专题命中 多模态RAG :RAG(title,abstract);retrieval-augmented generation(abstract)

AI总结 RAPTOR-AI通过分层多模态RAG框架,结合经验驱动的代理控制和LoRA适应,提升灾害响应中的检索精度、情境基础性和任务分解准确性。

Comments 8 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏