arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 8633 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 557 篇

2508.10337 2026-01-15 cs.AI cs.LG 79%

A Curriculum Learning Approach to Reinforcement Learning: Leveraging RAG for Multimodal Question Answering

一种基于RAG的强化学习课程学习方法:用于多模态问答

Chenliang Zhang, Lin Wang, Yuanyuan Lu, Yusheng Qi, Kexin Wang, Peixu Hou, Wenshi Chen

机构 * Meituan(美团)

专题命中 多模态RAG :RAG(title);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本文提出了一种结合课程学习与强化学习的方法,用于多模态问答任务,通过检索增强生成系统在挑战中取得优异成绩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08226 2026-01-14 cs.CV cs.AI 79%

Knowledge-based learning in Text-RAG and Image-RAG

基于知识的学习在Text-RAG和Image-RAG中的应用

Alexander Shim, Khalil Saieh, Samuel Clarke

机构 * Florida International University(佛罗里达国际大学)

专题命中 多模态RAG :RAG(title,abstract);分类 cs.AI

AI总结 本研究通过比较基于文本和图像的RAG方法,探讨了如何利用外部知识减少幻觉问题并提升胸部X光图像疾病检测的准确性。

Comments 9 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18987 2025-12-23 cs.RO cs.CL cs.CV 79%

Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation

语义可感知的多模态检索:基于具身记忆的层次化移动操作

Ryosuke Korekata, Quanting Xie, Yonatan Bisk, Komei Sugiura

机构 * Keio University(keio大学) Keio AI Research Center(keio人工智能研究中心) Carnegie Mellon University(卡内基梅隆大学)

专题命中 多模态RAG :RAG(title,abstract);分类 cs.CL

AI总结 本研究提出Affordance RAG框架,通过构建具有可操作性的具身记忆,提升机器人在开放词汇移动操作中的检索性能和任务成功率。

Comments Accepted to IEEE RA-L, with presentation at ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21002 2025-11-27 cs.CV cs.AI 79%

Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning

知识完善视觉:一种多模态实体感知检索增强生成框架用于新闻图像描述

Xiaoxing You, Qiang Huang, Lingyu Li, Chi Zhang, Xiaopeng Liu, Min Zhang, Jun Yu

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);分类 cs.AI

AI总结 MERGE提出了一种多模态实体感知检索增强生成框架,通过构建实体中心知识库和改进跨模态对齐,提升新闻图像描述质量和命名实体识别性能。

Comments Accepted to AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09266 2025-10-13 cs.CL 79%

CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation

Kaiwen Wei, Xiao Liu, Jie Zhang, Zijian Wang, Ruida Liu, Yuming Yang, Xin Xiao, Xiao Sun, Haoyang Zeng, Changzai Pan, Yidan Zhang, Jiang Zhong, Peijin Wang, Yingchao Feng

机构 * Chongqing University(重庆大学) Independent Researcher(独立研究者) University of the Chinese Academy of Sciences(中国科学院大学) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航天信息研究所)

专题命中 多模态RAG :retrieval-augmented generation(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11937 2025-09-16 cs.SE cs.AI 79%

MMORE: Massive Multimodal Open RAG & Extraction

Alexandre Sallinen, Stefan Krsteski, Paul Teiletche, Marc-Antoine Allard, Baptiste Lecoeur, Michael Zhang, Fabrice Nemo, David Kalajdzic, Matthias Meyer, Mary-Anne Hartley

机构 * École Polytechnique Fédérale de Lausanne (EPFL), Switzerland(瑞士联邦理工学院洛桑校区) ETH Zürich, Switzerland(瑞士苏黎世联邦理工学院) T.H. Chan School of Public Health, Harvard University, USA(哈佛大学T.H. Chan公共卫生学院)

专题命中 多模态RAG :RAG(title,abstract);分类 cs.AI

Comments This paper was originally submitted to the CODEML workshop for ICML 2025. 9 pages (including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14864 2025-02-21 cs.AI cs.CV 79%

Benchmarking Multimodal RAG through a Chart-based Document Question-Answering Generation Framework

Yuming Yang, Jiang Zhong, Li Jin, Jingwang Huang, Jingpeng Gao, Qing Liu, Yang Bai, Jingyuan Zhang, Rui Jiang, Kaiwen Wei

专题命中 多模态RAG :RAG(title);retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10834 2025-01-22 cs.CV cs.AI cs.LG 79%

Visual RAG: Expanding MLLM visual knowledge without fine-tuning

Mirco Bonomo, Simone Bianco

专题命中 多模态RAG :RAG(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06962 2025-06-17 cs.CV 79%

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation

Jingyuan Qi, Zhiyang Xu, Qifan Wang, Lifu Huang

机构 * Virginia Tech(弗吉尼亚理工大学) Meta UC Davis(加州大学戴维斯分校)

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(comments)

Comments Image Generation, Retrieval Augmented Generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21529 2026-08-25 cs.CV cs.CL cs.IR 新提交 79%

DamageScope: Vision-Language Retrieval at Scale for Disaster Damage Assessment from Satellite Imagery

DamageScope:面向卫星图像灾害损失评估的大规模视觉-语言检索方法

Ravi K. Rajendran, Biplob Debnath, Murugan Sankaradas, Srimat T. Chakradhar

机构 * NEC Laboratories America(美国 NEC 实验室)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.CL

AI总结 DamageScope是一种结合VLMs与LLMs的检索增强框架,通过多向量嵌入聚类和双存储架构,实现了卫星图像灾害损失评估的高效自动化,可扩展性与运营效率优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13706 2026-08-19 cs.CL cs.AI 版本更新 79%

CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA

CLAIR-Fin:用于跨模态金融问答中声明级验证与自适应辩论的对抗性多智能体框架

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Jubayer Al Mahmud, M. F. Mridha, Md. Alam Hossain

机构 * Ahsanullah University of Science and Technology(阿萨努拉科技大学) Jashore University of Science and Technology(杰索尔科技大学) American International University - Bangladesh(孟加拉国美国国际大学)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出CLAIR-Fin九智能体框架,针对跨模态金融问答的声明级验证与自适应辩论,在BB-FinQA-X数据集上提升了模型忠实度,且弃权比例合理,优于相关基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03836 2026-07-07 cs.CV cs.AI cs.CL 新提交 79%

When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts

何时越简单越好:评估中世纪拉丁文手稿的翻译管道

Nguyen Kim Hai Bui, Md. Easin Arafat, Tamás Gábor Orosz, Mufti Mahmud

机构 * Eötvös Loránd University(厄特沃什·罗兰大学) King Fahd University of Petroleum and Minerals(法赫德国王石油与矿产大学)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 针对历史手稿翻译难题,提出评估中世纪拉丁文手稿图像到翻译流程的框架。通过CATMuS数据集对比发现领域特定OCR模型优势,引入新数据集IPC,实验揭示简单管道表现更佳,为低资源历史场景翻译系统部署提供基准与指导。

Comments 17 pages, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28607 2026-05-28 cs.AI cs.CL 79%

Adaptive Multimodal Agents-Based Framework for Automatic Workflow Execution

基于自适应多智能体框架的自动工作流执行

Susanna Cifani, Mario Luca Bernardi, Marta Cimitile

机构 * Sapienza University of Rome(罗马萨皮恩扎大学) Department of Engineering University of Sannio(萨尼奥大学工程系) Faculty of Jurisprudence Unitelma Sapienza University(法理学院萨皮恩扎大学)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 提出一种多模态多智能体框架,通过离线构建拓扑知识库和在线自适应检索增强生成与闭环协作验证,实现自动工作流执行。

Comments Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses. Accepted for publication at the 2026 IEEE International Conference on Evolving and Adaptive Intelligent Systems (EAIS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18774 2026-05-20 cs.IR cs.AI 79%

M3DocDep: Multi-modal, Multi-page, Multi-document Dependency Chunking with Large Vision-Language Models

M3DocDep: 多模态、多页、多文档依赖分块方法基于大视觉-语言模型

Joongmin Shin, Jeongbae Park, Jaehyung Seo, Heuiseok Lim

机构 * Human-inspired AI Research, Korea University(韩国大学人智AI研究所) Computer Science and Engineering, Konkuk University(konkuk大学计算机科学与工程系) Department of Computer Science and Engineering, Korea University(韩国大学计算机科学与工程系)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 本文提出M3DocDep,一种基于大视觉-语言模型的多模态、多页、多文档依赖分块方法,通过恢复块级依赖并构建分块,提高了长多页多模态文档的检索和问答质量。

Comments Accepted to CVPR2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16275 2026-05-19 cs.CY cs.AI cs.CL cs.MM 79%

AI Slop or AI-enhancement? Student perceptions of AI-generated media for an English for Academic Purposes course

AI 产出物还是AI增强?英语学术用途课程中学生对AI生成媒体的看法

David James Woo, Deliang Wang, Kai Guo

机构 * Everwrite Limited(Everwrite有限公司) Faculty of Education, The University of Hong Kong(香港大学教育学院) Faculty of Education, The Chinese University of Hong Kong(香港中文大学教育学院)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨了AI生成内容在EAP课程中的教学效果,通过混合方法分析发现学生偏好视觉化内容,视频与学业表现正相关,但高认知负荷与成绩负相关,表明需合理设计内容以提升学习效果。

Comments 23 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11864 2026-05-13 cs.IR cs.AI cs.CV cs.MM 79%

Very Efficient Listwise Multimodal Reranking for Long Documents

非常高效的长文档多模态重排序方法

Yiqun Sun, Pengfei Wei, Lawrence B. Hsieh

机构 * Magellan Technology Research Institute (MTRI)(马杰拉技术研究院(MTRI))

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 本文提出ZipRerank,通过轻量级查询-图像早期交互机制和单次前向传递消除自回归解码,实现高效多模态重排序,实验表明其在MMDocIR基准上性能优异且显著降低LLM推理延迟。

Comments To appear in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07241 2026-07-10 cs.SD 新提交 78%

Rag Classification of Tagore Songs using Symbolic Music Notation and Novel Weighted Distance Measures

使用符号音乐记谱法和新型加权距离度量对泰戈尔歌曲进行拉格分类

Chandan Misra, Swarup Chattopadhyay

机构 * XIM University(西姆大学)

专题命中 多模态RAG :RAG(title,abstract)

AI总结 该研究针对罗宾德拉·桑吉特歌曲拉格识别难题,将其转化为监督分类问题,利用符号乐谱记谱法构建数据集,探讨多种距离度量,引入加权欧几里得距离,在k近邻框架下改进拉格分类,更好捕捉旋律特征。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20414 2026-05-26 eess.AS 78%

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding

PlanRAG-Audio:面向长音频理解的规划与检索增强生成

Masao Someki, Chien-yu Huang, Siddhant Arora, Samuele Cornell, Markus Müller, Nathan Susanj, Rupak V Swaminathan, Grant P Strimel, Jing Liu, Shinji Watanabe

专题命中 多模态RAG :retrieval augmented generation(title);retrieval-augmented generation(abstract)

AI总结 提出PlanRAG-Audio框架,通过规划查询所需模态和时间跨度并仅检索相关信息,实现长音频的高效推理,提升准确率并稳定性能。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04372 2026-04-07 cs.CV 78%

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning

图到帧RAG:面向无训练和可审计的视频推理的视觉空间知识融合

Songyuan Yang, Weijiang Yu, Ziyu Liu, Guijian Tang, Wenjing Yang, Huibin Tan, Nong Xiao

机构 * National University of Defense Technology(国防科技大学) Sun Yat-sen University(中山大学)

专题命中 多模态RAG :RAG(title,abstract)

AI总结 本文提出G2F-RAG,通过视觉空间知识融合提升视频推理的可解释性和效率,减少认知负担并保留可追溯的证据轨迹。

Comments Accepted at CVPR 2026. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02258 2026-03-24 cs.CV 78%

Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning

病理代理RAG:通过强化学习实现多模态代理检索增强生成用于病理学视觉语言模型

Wenchuan Zhang, Jingru Guo, Hengzhe Zhang, Penghao Zhang, Jie Chen, Shuwan Zhang, Zhang Zhang, Yuhao Yi, Hong Bu

专题命中 多模态RAG :retrieval-augmented generation(title);RAG(abstract)

AI总结 本文提出Patho-AgenticRAG,通过强化学习实现多模态代理检索增强生成,解决病理学视觉语言模型在高分辨率、复杂组织结构和临床语义上的挑战,提升诊断准确性。

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(35): 29921-29929, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23483 2025-12-30 cs.CV 78%

TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding

TV-RAG:一种具有时间意识和语义熵权的长视频检索与理解框架

Zongsheng Cao, Yangfan He, Anran Liu, Feng Chen, Zepeng Wang, Jun Xie

机构 * Researcher(研究者)

专题命中 多模态RAG :RAG(title,abstract)

AI总结 TV-RAG通过时间衰减检索和熵加权关键帧采样,提升长视频检索与理解性能,无需重新训练即可集成至现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05872 2025-12-10 cs.CV 78%

Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection

领域-RAG:跨领域少样本目标检测的检索引导组合图像生成

Yu Li, Xingyu Qiu, Yuqian Fu, Jie Chen, Tianwen Qian, Xu Zheng, Danda Pani Paudel, Yanwei Fu, Xuanjing Huang, Luc Van Gool, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Fuzhou University(福州大学) East China Normal University(华东师范大学) HKUST(GZ)(香港科技大学(广州))

专题命中 多模态RAG :RAG(title,abstract)

AI总结 Domain-RAG通过检索引导的组合图像生成方法,解决跨领域少样本目标检测中的领域对齐与背景生成问题,实现高质量样本生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07538 2025-09-11 cs.CV 78%

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

Peijin Xie, Shun Qian, Bingquan Liu, Dexin Wang, Lin Sun, Xiangzheng Zhang

机构 * IEEE

专题命中 多模态RAG :RAG(title,abstract)

Comments 5 pages, 4 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00751 2025-09-03 cs.CV 78%

EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions

Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science, VNU-HCM(越南国家大学科学学院)

专题命中 多模态RAG :retriever(title,abstract)

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06496 2025-08-12 cs.CV cs.MA 78%

Med-GRIM: Enhanced Zero-Shot Medical VQA using prompt-embedded Multimodal Graph RAG

Rakesh Raj Madavan, Akshat Kaimal, Hashim Faisal, Chandrakala S

机构 * Shiv Nadar University Chennai(施瓦斯纳大学钦奈)

专题命中 多模态RAG :RAG(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08576 2025-03-12 cs.CV 78%

RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding

Xichen Tan, Yunfan Ye, Yuanjing Luo, Qian Wan, Fang Liu, Zhiping Cai

专题命中 多模态RAG :RAG(title,abstract)

Comments 37 pages, 36 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03340 2024-08-08 cs.MM cs.CV 78%

An Empirical Comparison of Video Frame Sampling Methods for Multi-Modal RAG Retrieval

Mahesh Kandhare, Thibault Gisselbrecht

专题命中 多模态RAG :RAG(title,abstract)

Comments 19 pages, 24 figures (65 images)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18386 2026-08-20 cs.CV cs.AI 新提交 77%

TTSD-FAR: Test-Time Self-Distillation with Fisher-Anchored Restoration for Missing-Modality Emotion Recognition in LVLMs

TTSD-FAR:用于大视频语言模型中缺失模态情感识别的带Fisher锚定恢复的测试时自蒸馏

Muhammad Haseeb Aslam, Alessandro Koerich, Marco Pedersoli, Ali Etemad, Eric Granger

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval augmented generation(abstract);分类 cs.AI

AI总结 针对大视频语言模型测试时的模态缺失问题,提出带Fisher锚定恢复的测试时自蒸馏框架,在多数据集模态缺失场景下性能优于基线方法且保持稳定。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25041 2026-07-29 cs.IR cs.CR cs.CV cs.LG 新提交 77%

ScoreShield: Differentially Private Release of Similarity Scores

ScoreShield:相似度分数的差分隐私发布

Behrooz Razeghi, Parsa Rahimi

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR

AI总结 研究生物识别等应用中相似度分数隐私保护问题,提出ScoreShield先扰动后投影机制,添加校准高斯噪声并投影到可行性集,满足(ε,δ)-DP,给出效用保证,在多任务中评估,改进了风险的n依赖性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01814 2026-07-03 cs.AI 新提交 77%

MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support

MMIR-TCM:面向中医临床决策支持的记忆集成多模态推理与检索

Lihui Luo, Joongwon Chae, Ziyan Chen, Yang Liu, Siyi Cheng, Weihan Gao, Zelin Zeng, Xiaoming Yin, Samaneh Beheshti Kashi, Dongmei Yu, Lian Zhang, Jing Sui, Zeming Liang, Jiansong Ji, Peter E. Lobie, Peiwu Qin

机构 * Institute of Biopharmaceutics and Health Engineering, Tsinghua Shenzhen International Graduate School(清华深圳国际研究生院生物医药与健康工程研究院) Chinese Medicine Guangdong Laboratory(广东省中医药实验室) Beijing Normal University(北京师范大学) First Hospital of Hebei Medical University(河北医科大学第一医院) Lishui Hospital of Zhejiang University(浙江大学丽水医院) XiaoMing TCM Hospital(小明中医医院)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 提出MMIR-TCM框架,结合多模态大模型、记忆增强分割与检索增强生成,解决中医舌诊中的主观性和语义鸿沟问题,在MedTCM数据集上优于GPT-4o等模型。

详情

展开后加载摘要…

URL PDF HTML 收藏