arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 551 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 551 篇

2510.06820 2026-02-24 cs.CV cs.LG 50%

Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking

大规模视觉-语言重排序的高效判别联合编码器

Mitchell Keren Taraday, Shahaf Wagner, Chaim Baskin

机构 * INSIGHT Lab, Ben-Gurion University of the Negev(本·古里安内盖夫大学INSIGHT实验室)

专题命中 多模态RAG :vector search(abstract)

AI总结 EDJE通过离线预计算和轻量级注意力适配器,实现了高效视觉-语言重排序,显著降低存储和计算成本,提升大规模检索性能。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16493 2026-02-19 cs.CV 50%

MMA: Multimodal Memory Agent

MMA:多模态记忆代理

Yihao Lu, Wanru Cheng, Zeyu Zhang, Hao Tang

机构 * School of Computer Science, Peking University(北京大学计算机学院)

专题命中 多模态RAG :RAG(abstract)

AI总结 MMA通过动态可靠性评分和冲突感知网络共识,提升多模态代理在长horizon任务中的记忆检索与决策能力,同时在多个基准测试中展现优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.14322 2026-02-16 cs.CV 50%

Spatio-Temporal driven Attention Graph Neural Network with Block Adjacency matrix (STAG-NN-BA) for Remote Land-use Change Detection

基于块邻接矩阵的时空驱动注意力图神经网络(STAG-NN-BA)用于遥感土地利用变化检测

Usman Nazir, Wadood Islam, Sara Khalid, Murtaza Taj

专题命中 多模态RAG :RAG(abstract)

AI总结 本文提出STAG-NN-BA模型,利用超像素和块邻接矩阵实现遥感土地利用变化检测,优于传统图和非图基线。

Journal ref AAAI Symposium 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10656 2026-02-12 eess.AS cs.SD 50%

AudioRAG: A Challenging Benchmark for Audio Reasoning and Information Retrieval

AudioRAG:一个用于音频推理和信息检索的挑战性基准

Jingru Lin, Chen Zhang, Tianrui Wang, Haizhou Li

专题命中 多模态RAG :retrieval-augmented generation(abstract)

AI总结 AudioRAG是一个用于评估音频推理和信息检索能力的挑战性基准,通过整合检索增强生成技术,为未来研究提供更强大的基准。

Comments Accepted by Audio-AAAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02537 2026-02-04 cs.CV cs.LG 50%

WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models

WorldVQA:评估多模态大语言模型的原子视觉世界知识

Runjie Zhou, Youbo Shao, Haoyu Lu, Bowei Xing, Tongtong Bai, Yujie Chen, Jie Zhao, Lin Sui, Haotian Yao, Zijia Zhao, Hao Yang, Haoning Wu, Zaida Zhou, Jinguo Zhu, Zhiqi Huang, Yiping Bao, Yangyang Liu, Y. Charles, Xinyu Zhou

机构 * Moonshot AI

专题命中 多模态RAG :knowledge retrieval(abstract)

AI总结 WorldVQA通过评估多模态大语言模型的原子视觉知识,建立衡量其事实性和百科全书广度的基准测试。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13801 2026-01-21 cs.RO 50%

HoverAI: An Embodied Aerial Agent for Natural Human-Drone Interaction

HoverAI: 一种用于自然人-无人机交互的具身空中代理

Yuhua Jin, Nikita Kuzmin, Georgii Demianchuk, Mariya Lezina, Fawad Mehboob, Issatay Tokmurziyev, Miguel Altamirano Cabrera, Muhammad Ahsan Mustafa, Dzmitry Tsetserukou

机构 * Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Skolkovo Institute of Science and Technology(斯克尔科沃信息科技研究所)

专题命中 多模态RAG :RAG(abstract)

AI总结 HoverAI通过结合无人机移动、视觉投影和对话式AI,实现了人-无人机自然交互的具身代理,提升了空间感知与社交响应能力。

Comments This paper has been accepted for publication at LBR HRI 2026 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11537 2026-01-21 cs.HC 50%

Building AI-based advisory services for smallholder farmers: Technical learnings from the AIEP Initiative

为小农户建设基于AI的咨询服务:AIEP计划的技术经验

Stewart Collis, Florence Kinyua, Vikram Kumar, Howard Lakougna, Christian Merz, Kirti Pandey, Christian Resch

专题命中 多模态RAG :RAG(abstract)

AI总结 AIEP计划通过AI技术为小农户提供农业咨询服务,发现多语言语音交互和语料库编纂是关键挑战,同时强调数据共享和评估基准的重要性。

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23044 2026-01-19 cs.CV 50%

Video-Browser: Towards Agentic Open-web Video Browsing

Video-Browser: 向具有代理能力的开放网页视频浏览迈进

Zhengyang Liang, Yan Shu, Xiangrui Liu, Minghao Qin, Kaixin Liang, Nicu Sebe, Zheng Liu, Lizi Liao

专题命中 多模态RAG :RAG(abstract)

AI总结 Video-Browser通过金字塔感知技术,在开放式网络视频浏览中实现37.5%的相对提升,同时减少58.3%的token消耗。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04540 2025-12-17 cs.CV 50%

VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management

VideoMem: 通过自适应内存管理增强超长视频理解

Hongbo Jin, Qingyuan Wang, Wenhao Zhang, Yang Liu, Sijie Cheng

机构 * School of Electronic and Computer Engineering, Peking University(电子与计算机工程学院,北京大学) Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学)

专题命中 多模态RAG :RAG(abstract)

AI总结 VideoMem通过自适应内存管理框架,有效提升超长视频理解任务的性能,采用PRPO算法和两个核心模块实现高效训练和长期记忆保留。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02395 2025-12-09 cs.CV 50%

Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch

Skywork-R1V4:通过图像与深度研究交织思考实现代理多模态智能

Yifan Zhang, Liang Hu, Haofeng Sun, Peiyu Wang, Yichen Wei, Shukang Yin, Jiangbo Pei, Wei Shen, Peng Xia, Yi Peng, Tianyidan Xie, Eric Li, Yang Liu, Xuchen Song, Yahui Zhou

机构 * Skywork AI

专题命中 多模态RAG :knowledge retrieval(abstract)

AI总结 Skywork-R1V4通过交织推理实现多模态代理智能,仅用监督学习在少数据上训练,超越现有模型在多个基准测试中的表现。

Comments 21 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18883 2025-11-24 cs.CV 50%

Universal Video Temporal Grounding with Generative Multi-modal Large Language Models

通用视频时间定位与生成多模态大语言模型

Zeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang, Yanfeng Wang, Weidi Xie

机构 * SAI, Shanghai Jiao Tong University(上海交通大学SAI实验室) ByteDance Seed(字节跳动种子)

专题命中 多模态RAG :retriever(abstract)

AI总结 本文提出UniTime模型,利用生成多模态大语言模型实现通用视频时间定位,有效处理多类型视频并提升VideoQA任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06752 2025-11-11 cs.CV 50%

Med-SORA: Symptom to Organ Reasoning in Abdomen CT Images

You-Kyoung Na, Yeong-Jun Cho

机构 * Chonnam National University(全南国立大学)

专题命中 多模态RAG :RAG(abstract)

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05199 2025-11-10 cs.RO 50%

Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation

Yichen Zhu, Feifei Feng

机构 * Midea Group, AI Research Center(美的集团人工智能研究中心)

专题命中 多模态RAG :retriever(abstract)

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26386 2025-10-29 cs.CV 50%

PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer

Zhiwei Yang, Chen Gao, Mike Zheng Shou

机构 * Xidian University(西安电子科技大学) Show Lab, National University of Singapore(新加坡国立大学Show实验室)

专题命中 多模态RAG :RAG(abstract)

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16510 2025-10-21 q-bio.BM 50%

CryoDyna: Multiscale end-to-end modeling of cryo-EM macromolecule dynamics with physics-aware neural network

Chengwei Zhang, Shimian Li, Yihao Niu, Zhen Zhu, Sihao Yuan, Sirui Liu, Yi Qin Gao

专题命中 多模态RAG :RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19203 2025-09-24 cs.CV 50%

Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions

Ioanna Ntinou, Alexandros Xenos, Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

机构 * Queen Mary University of London(伦敦女王大学) Samsung AI Centre(三星人工智能中心) Technical University of Iași(伊阿苏技术大学)

专题命中 多模态RAG :retriever(abstract)

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01960 2025-09-23 cs.LG 50%

MPIC: Position-Independent Multimodal Context Caching System for Efficient MLLM Serving

Shiju Zhao, Junhao Hu, Rongxiao Huang, Jiaqi Zheng, Guihai Chen

机构 * State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) Nanjing University(南京大学) School of Computer Science(计算机学院)

专题命中 多模态RAG :retrieval-augmented generation(abstract)

Comments 17 pages, 13 figures, the second version

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13899 2025-09-18 cs.HC 50%

AI as a teaching tool and learning partner

Steven Watterson, Sarah Atkinson, Elaine Murray, Andrew McDowell

专题命中 多模态RAG :RAG(abstract)

Comments 6 Pages, 1 Figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01970 2025-08-29 cs.LG 50%

Improving Hospital Risk Prediction with Knowledge-Augmented Multimodal EHR Modeling

Rituparna Datta, Jiaming Cui, Zihan Guan, Vishal G. Reddy, Joshua C. Eby, Gregory Madden, Rupesh Silwal, Anil Vullikanti

机构 * Department of Computer Science, University of Virginia(大学计算机科学系) University of Virginia School of Medicine(弗吉尼亚大学医学院) Virginia Polytechnic Institute and State University(弗吉尼亚理工学院和州立大学) Biocomplexity Institute and Initiative, University of Virginia(大学生物复杂性研究所) Division of Infectious Diseases & International Health, University of Virginia School of Medicine(大学感染性疾病与国际卫生分会)

专题命中 多模态RAG :knowledge retrieval(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11974 2025-07-17 cs.RO 50%

A Review of Generative AI in Aquaculture: Foundations, Applications, and Future Directions for Smart and Sustainable Farming

Waseem Akram, Muhayy Ud Din, Lyes Saad Soud, Irfan Hussain

机构 * Khalifa University Center for Autonomous Robotic Systems (KUCARS), Khalifa University, United Arab Emirates(卡里法大学自主机器人系统中心(KUCARS)、卡里法大学、阿拉伯联合酋长国)

专题命中 多模态RAG :retrieval augmented generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02900 2025-07-11 cs.CV 50%

MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine

Yunfei Xie, Ce Zhou, Lang Gao, Juncheng Wu, Xianhang Li, Hong-Yu Zhou, Sheng Liu, Lei Xing, James Zou, Cihang Xie, Yuyin Zhou

机构 * Huazhong University of Science and Technology(华中科技大学) UC Santa Cruz(加州大学圣克ruz分校) Harvard University(哈佛大学) Stanford University(斯坦福大学)

专题命中 多模态RAG :retrieval-augmented generation(abstract)

Comments The dataset is publicly available at https://yunfeixie233.github.io/MedTrinity-25M/. Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18899 2025-06-24 cs.CV 50%

FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation

Kaiyi Huang, Yukun Huang, Xintao Wang, Zinan Lin, Xuefei Ning, Pengfei Wan, Di Zhang, Yu Wang, Xihui Liu

机构 * The University of Hong Kong(香港大学) Kuaishou Technology(快手科技) Microsoft Research(微软研究院) Tsinghua University(清华大学)

专题命中 多模态RAG :RAG(abstract)

Comments Project Page: https://filmaster-ai.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17837 2025-06-24 cs.CV 50%

Time-Contrastive Pretraining for In-Context Image and Video Segmentation

Assefa Wahd, Jacob Jaremko, Abhilash Hareendranathan

机构 * Department of Radiology(放射科部门)

专题命中 多模态RAG :retriever(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10533 2025-05-16 cs.CV cs.LG 50%

Enhancing Multi-Image Question Answering via Submodular Subset Selection

Aaryan Sharma, Shivansh Gupta, Samar Agarwal, Vishak Prasad C., Ganesh Ramakrishnan

机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔学院)

专题命中 多模态RAG :retriever(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05782 2025-03-18 cs.SD cs.CV cs.LG cs.MM eess.AS 50%

Sequential Contrastive Audio-Visual Learning

Ioannis Tsiamas, Santiago Pascual, Chunghsin Yeh, Joan Serrà

专题命中 多模态RAG :hybrid retrieval(abstract)

Comments ICASSP 2025. Version 1 contains more details

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02692 2025-03-05 cs.CE econ.GN q-fin.EC 50%

FinArena: A Human-Agent Collaboration Framework for Financial Market Analysis and Forecasting

Congluo Xu, Zhaobin Liu, Ziyang Li

专题命中 多模态RAG :RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21068 2025-03-03 cs.SE 50%

GUIDE: LLM-Driven GUI Generation Decomposition for Automated Prototyping

Kristian Kolthoff, Felix Kretzer, Christian Bartelt, Alexander Maedche, Simone Paolo Ponzetto

专题命中 多模态RAG :retrieval-augmented generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17038 2025-02-25 cs.MM 50%

Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction

Jiacheng Lu, Mingyuan Xiao, Weijian Wang, Yuxin Du, Zhengze Wu, Cheng Hua

专题命中 多模态RAG :retriever(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20725 2024-12-31 cs.CV 50%

Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling

Min Zhang, Zilin Wang, Liyan Chen, Kunhong Liu, Juncong Lin

专题命中 多模态RAG :retrieval-augmented generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09936 2024-12-16 cs.CV 50%

CaLoRAify: Calorie Estimation with Visual-Text Pairing and LoRA-Driven Visual Language Models

Dongyu Yao, Keling Yao, Junhong Zhou, Yinghao Zhang

专题命中 多模态RAG :RAG(abstract)

Comments Disclaimer: This work is part of a course project and reflects ongoing exploration in the field of vision-language models and calorie estimation. Findings and conclusions are subject to further validation and refinement

详情

展开后加载摘要…

URL PDF HTML 收藏