arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 551 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 551 篇

2505.10610 2025-10-07 cs.CV cs.CL 57%

MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

Zhaowei Wang, Wenhao Yu, Xiyu Ren, Jipeng Zhang, Yu Zhao, Rohit Saxena, Liang Cheng, Ginny Wong, Simon See, Pasquale Minervini, Yangqiu Song, Mark Steedman

机构 * CSE Department, HKUST(香港科技大学计算机科学与工程系) Tencent AI Seattle Lab(腾讯AI西雅图实验室) University of Edinburgh(爱丁堡大学) NVIDIA AI Technology Center (NVAITC), NVIDIA, Santa Clara, USA(英伟达圣克拉拉人工智能技术中心)

专题命中 多模态RAG :RAG(abstract);分类 cs.CL

Comments Accepted as a spotlight at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00088 2025-10-02 cs.AI cs.CY 57%

Judging by Appearances? Auditing and Intervening Vision-Language Models for Bail Prediction

Sagnik Basu, Shubham Prakash, Ashish Maruti Barge, Siddharth D Jaiswal, Abhisek Dash, Saptarshi Ghosh, Animesh Mukherjee

机构 * Indian Institute of Technology Kharagpur(印度理工学院Khargapur分校) Max Planck Institute for Software Systems(马克斯·普朗克软件系统研究所)

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24350 2025-09-30 cs.CV cs.AI 57%

Dynamic Orchestration of Multi-Agent System for Real-World Multi-Image Agricultural VQA

Yan Ke, Xin Yu, Heming Du, Scott Chapman, Helen Huang

机构 * The University of Queensland(昆士兰大学)

专题命中 多模态RAG :retriever(abstract);分类 cs.AI

Comments 13 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05748 2025-09-03 cs.IR 57%

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Xinyu Geng, Peng Xia, Zhen Zhang, Xinyu Wang, Qiuchen Wang, Ruixue Ding, Chenxi Wang, Jialong Wu, Yida Zhao, Kuan Li, Yong Jiang, Pengjun Xie, Fei Huang, Jingren Zhou

专题命中 多模态RAG :RAG(abstract);分类 cs.IR

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12263 2025-09-01 cs.CV cs.AI 57%

Region-Level Context-Aware Multimodal Understanding

Hongliang Wei, Xianqi Zhang, Xingtao Wang, Xiaopeng Fan, Debin Zhao

机构 * Faculty of Computing, Harbin Institute of Technology(计算机学院,哈尔滨工业大学) Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究院长) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国)

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19319 2025-08-28 eess.IV cs.AI cs.CV 57%

MedVQA-TREE: A Multimodal Reasoning and Retrieval Framework for Sarcopenia Prediction

Pardis Moradbeiki, Nasser Ghadiri, Sayed Jalal Zahabi, Uffe Kock Wiil, Kristoffer Kittelmann Brockhattingen, Ali Ebrahimi

机构 * Department of Electrical and Computer Engineering, Isfahan University of Technology(电气与计算机工程系,伊斯法罕技术大学) SDU Health Informatics and Technology, The Maersk Mc-Kinney Moller Institute, University of Southern Denmark(南部丹麦大学健康信息学与技术,马士基麦金尼莫勒研究所) Geriatric Research Unit, Department of Clinical Research, University of Southern Denmark(老年医学研究单元,临床研究系,南部丹麦大学)

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18108 2025-08-26 cs.CL 57%

SentiMM: A Multimodal Multi-Agent Framework for Sentiment Analysis in Social Media

Xilai Xu, Zilin Zhao, Chengye Song, Zining Wang, Jinhe Qiang, Jiongrui Yan, Yuhuai Lin

机构 * College of Information and Electrical Engineering, China Agricultural University(信息与电气工程学院,中国农业大学) College of Software, Jilin University(软件学院,吉林大学) College of Communication Engineering, Jilin University(通信工程学院,吉林大学)

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06328 2025-08-11 cs.IR 57%

M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation

Zhiyou Xiao, Qinhan Yu, Binghui Li, Geng Chen, Chong Chen, Wentao Zhang

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.IR

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12884 2025-07-01 cs.LG cs.AI cs.CV 57%

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks

Yuanze Hu, Zhaoxin Fan, Xinyu Wang, Gen Li, Ye Qiu, Zhichao Yang, Wenjun Wu, Kejian Wu, Yifan Sun, Xiaotie Deng, Jin Dong

机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(北京未来区块链与隐私计算先进创新中心) Beihang University(北京航空航天大学) Hangzhou International Innovation Institute(杭州国际创新研究院) Xreal Renmin University(中国人民大学) Peking University(北京大学) Beijing Academy of Blockchain and Edge Computing (BABEC)(北京区块链与边缘计算研究院)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12831 2025-07-01 eess.IV cs.AI cs.CV 57%

Segment as You Wish -- Free-Form Language-Based Segmentation for Medical Images

Longchao Da, Rui Wang, Xiaojian Xu, Parminder Bhatia, Taha Kass-Hout, Hua Wei, Cao Xiao

机构 * Arizona State University(亚利桑那州立大学)

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 19 pages, 9 as main content. The paper was accepted to KDD2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10756 2025-06-13 cs.RO cs.AI 57%

Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding

Yuhang Zhang, Haosheng Yu, Jiaping Xiao, Mir Feroskhan

机构 * School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院,南洋理工大学)

专题命中 多模态RAG :retriever(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10913 2025-06-11 cs.SD cs.AI eess.AS 57%

Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning

Choi Changin, Lim Sungjun, Rhee Wonjong

机构 * Interdisciplinary Program in Artificial Intelligence(人工智能交叉学科项目) Department of Intelligence and Information(智能与信息系)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02470 2025-06-04 cs.AI 57%

A Smart Multimodal Healthcare Copilot with Powerful LLM Reasoning

Xuejiao Zhao, Siyan Liu, Su-Yin Yang, Chunyan Miao

机构 * Joint NTU-UBC Research Centre of Excellence in Active Living for the Elderly (LILY), NTU(联合NTU-UBC老龄化积极生活卓越研究中心(LILY),NTU) College of Computing and Data Science, Nanyang Technological University (NTU), Singapore(计算与数据科学学院,南洋理工大学(NTU),新加坡) Tan Tock Seng Hospital, Singapore(坦 tok sing 医院,新加坡) Woodlands Health, Singapore(伍德兰兹健康,新加坡)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14318 2025-06-03 cs.CV cs.CL 57%

RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection

Wenjun Hou, Yi Cheng, Kaishuai Xu, Heng Li, Yan Hu, Wenjie Li, Jiang Liu

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Research Institute of Trustworthy Autonomous Systems and Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学可信自主系统研究院和计算机科学与工程系) School of Computer Science, University of Nottingham Ningbo China(宁波大学计算机学院)

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.CL

Comments Accepted to ACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02466 2025-05-06 cs.IR 57%

Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality

Xueguang Ma, Luyu Gao, Shengyao Zhuang, Jiaqi Samantha Zhan, Jamie Callan, Jimmy Lin

专题命中 多模态RAG :retriever(abstract);分类 cs.IR

Comments Accepted in SIGIR 2025 (Demo)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16723 2025-04-24 cs.CV cs.AI 57%

Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering

Ali Anaissi, Junaid Akram, Kunal Chaturvedi, Ali Braytee

机构 * The University of Sydney, School of Computer Science(悉尼大学计算机科学学院) University of Technology Sydney, School of Computer Science(新南威尔士大学技术学院) University of Technology Sydney, TD School(新南威尔士大学TD学院) Australian Catholic University, Peter Faber Business School(澳大利亚天主教大学彼得·法伯商学院)

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 13 pages, 2 figures, 2025 International Conference on Computational Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00750 2025-04-17 cs.AI 57%

Beyond Text: Implementing Multimodal Large Language Model-Powered Multi-Agent Systems Using a No-Code Platform

Cheonsu Jeong

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 22 pages, 27 figures

Journal ref 2025 Journal of Intelligence and Information Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03241 2025-04-07 cs.CV cs.AI cs.LG 57%

Rotation Invariance in Floor Plan Digitization using Zernike Moments

Marius Graumann, Jan Marius Stürmer, Tobias Koch

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12836 2025-03-05 cs.CV cs.AI cs.LG 57%

AI-based association analysis for medical imaging using latent-space geometric confounder correction

Xianjing Liu, Bo Li, Meike W. Vernooij, Eppo B. Wolvius, Gennady V. Roshchupkin, Esther E. Bron

专题命中 多模态RAG :vector search(abstract);分类 cs.AI

Comments Accepted by Medical Image Analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05846 2025-03-04 cs.CV cs.CL 57%

Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines

Michael Toker, Hadas Orgad, Mor Ventura, Dana Arad, Yonatan Belinkov

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.CL

Comments Published in: ACL 2024 Project webpage: tokeron.github.io/DiffusionLensWeb

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08250 2025-02-24 cs.HC cs.AI 57%

OmniQuery: Contextually Augmenting Captured Multimodal Memory to Enable Personal Question Answering

Jiahao Nick Li, Zhuohao Jerry Zhang, Jiaju Ma

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments Paper accepted to the 2025 CHI Conference on Human Factors in Computing Systems (CHI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07365 2025-02-18 cs.IR cs.LG 57%

Multimodal semantic retrieval for product search

Dong Liu, Esther Lopez Ramos

专题命中 多模态RAG :dense retrieval(abstract);分类 cs.IR

Comments Accepted at EReL@MIR WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06781 2025-01-27 cs.AI 57%

Eliza: A Web3 friendly AI Agent Operating System

Shaw Walters, Sam Gao, Shakker Nerd, Feng Da, Warren Williams, Ting-Chien Meng, Amie Chow, Hunter Han, Frank He, Allen Zhang, Ming Wu, Timothy Shen, Maxwell Hu, Jerry Yan

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00846 2024-12-03 cs.AI 57%

Improving Multimodal LLMs Ability In Geometry Problem Solving, Reasoning, And Multistep Scoring

Avinash Anand, Raj Jaiswal, Abhishek Dharmadhikari, Atharva Marathe, Harsh Parimal Popat, Harshil Mital, Kritarth Prasad, Rajiv Ratn Shah, Roger Zimmermann

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16592 2024-10-23 cs.LG cs.CL cs.CY 57%

ViMGuard: A Novel Multi-Modal System for Video Misinformation Guarding

Andrew Kan, Christopher Kan, Zaid Nabulsi

专题命中 多模态RAG :retrieval augmented generation(abstract);分类 cs.CL

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13510 2024-10-18 cs.CL cs.CV 57%

GeoCoder: Solving Geometry Problems by Generating Modular Code through Vision-Language Models

Aditya Sharma, Aman Dalmia, Mehran Kazemi, Amal Zouaq, Christopher J. Pal

专题命中 多模态RAG :RAG(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18202 2024-10-01 cs.AI cs.MM 57%

WorldGPT: Empowering LLM as Multimodal World Model

Zhiqi Ge, Hongzhe Huang, Mingze Zhou, Juncheng Li, Guoming Wang, Siliang Tang, Yueting Zhuang

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.AI

Comments update v2

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11281 2024-09-18 cs.IR 57%

Beyond Relevance: Improving User Engagement by Personalization for Short-Video Search

Wentian Bao, Hu Liu, Kai Zheng, Chao Zhang, Shunyu Zhang, Enyun Yu, Wenwu Ou, Yang Song

专题命中 多模态RAG :dense retrieval(abstract);分类 cs.IR

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06450 2024-09-11 cs.RO cs.AI cs.ET 57%

Multimodal Large Language Model Driven Scenario Testing for Autonomous Vehicles

Qiujing Lu, Xuanhan Wang, Yiwei Jiang, Guangming Zhao, Mingyue Ma, Shuo Feng

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14723 2024-08-28 cs.CV cs.IR 57%

Snap and Diagnose: An Advanced Multimodal Retrieval System for Identifying Plant Diseases in the Wild

Tianqi Wei, Zhi Chen, Xin Yu

专题命中 多模态RAG :retriever(abstract);分类 cs.IR

详情

展开后加载摘要…

URL PDF HTML 收藏