arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 8633 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 557 篇

2410.10913 2025-06-11 cs.SD cs.AI eess.AS 57%

Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning

Choi Changin, Lim Sungjun, Rhee Wonjong

机构 * Interdisciplinary Program in Artificial Intelligence(人工智能交叉学科项目) Department of Intelligence and Information(智能与信息系)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02470 2025-06-04 cs.AI 57%

A Smart Multimodal Healthcare Copilot with Powerful LLM Reasoning

Xuejiao Zhao, Siyan Liu, Su-Yin Yang, Chunyan Miao

机构 * Joint NTU-UBC Research Centre of Excellence in Active Living for the Elderly (LILY), NTU(联合NTU-UBC老龄化积极生活卓越研究中心(LILY),NTU) College of Computing and Data Science, Nanyang Technological University (NTU), Singapore(计算与数据科学学院,南洋理工大学(NTU),新加坡) Tan Tock Seng Hospital, Singapore(坦 tok sing 医院,新加坡) Woodlands Health, Singapore(伍德兰兹健康,新加坡)

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14318 2025-06-03 cs.CV cs.CL 57%

RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection

Wenjun Hou, Yi Cheng, Kaishuai Xu, Heng Li, Yan Hu, Wenjie Li, Jiang Liu

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Research Institute of Trustworthy Autonomous Systems and Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学可信自主系统研究院和计算机科学与工程系) School of Computer Science, University of Nottingham Ningbo China(宁波大学计算机学院)

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.CL

Comments Accepted to ACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02466 2025-05-06 cs.IR 57%

Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality

Xueguang Ma, Luyu Gao, Shengyao Zhuang, Jiaqi Samantha Zhan, Jamie Callan, Jimmy Lin

专题命中 多模态RAG :retriever(abstract);分类 cs.IR

Comments Accepted in SIGIR 2025 (Demo)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16723 2025-04-24 cs.CV cs.AI 57%

Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering

Ali Anaissi, Junaid Akram, Kunal Chaturvedi, Ali Braytee

机构 * The University of Sydney, School of Computer Science(悉尼大学计算机科学学院) University of Technology Sydney, School of Computer Science(新南威尔士大学技术学院) University of Technology Sydney, TD School(新南威尔士大学TD学院) Australian Catholic University, Peter Faber Business School(澳大利亚天主教大学彼得·法伯商学院)

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 13 pages, 2 figures, 2025 International Conference on Computational Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00750 2025-04-17 cs.AI 57%

Beyond Text: Implementing Multimodal Large Language Model-Powered Multi-Agent Systems Using a No-Code Platform

Cheonsu Jeong

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 22 pages, 27 figures

Journal ref 2025 Journal of Intelligence and Information Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03241 2025-04-07 cs.CV cs.AI cs.LG 57%

Rotation Invariance in Floor Plan Digitization using Zernike Moments

Marius Graumann, Jan Marius Stürmer, Tobias Koch

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.12836 2025-03-05 cs.CV cs.AI cs.LG 57%

AI-based association analysis for medical imaging using latent-space geometric confounder correction

Xianjing Liu, Bo Li, Meike W. Vernooij, Eppo B. Wolvius, Gennady V. Roshchupkin, Esther E. Bron

专题命中 多模态RAG :vector search(abstract);分类 cs.AI

Comments Accepted by Medical Image Analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05846 2025-03-04 cs.CV cs.CL 57%

Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines

Michael Toker, Hadas Orgad, Mor Ventura, Dana Arad, Yonatan Belinkov

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.CL

Comments Published in: ACL 2024 Project webpage: tokeron.github.io/DiffusionLensWeb

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08250 2025-02-24 cs.HC cs.AI 57%

OmniQuery: Contextually Augmenting Captured Multimodal Memory to Enable Personal Question Answering

Jiahao Nick Li, Zhuohao Jerry Zhang, Jiaju Ma

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments Paper accepted to the 2025 CHI Conference on Human Factors in Computing Systems (CHI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07365 2025-02-18 cs.IR cs.LG 57%

Multimodal semantic retrieval for product search

Dong Liu, Esther Lopez Ramos

专题命中 多模态RAG :dense retrieval(abstract);分类 cs.IR

Comments Accepted at EReL@MIR WWW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06781 2025-01-27 cs.AI 57%

Eliza: A Web3 friendly AI Agent Operating System

Shaw Walters, Sam Gao, Shakker Nerd, Feng Da, Warren Williams, Ting-Chien Meng, Amie Chow, Hunter Han, Frank He, Allen Zhang, Ming Wu, Timothy Shen, Maxwell Hu, Jerry Yan

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00846 2024-12-03 cs.AI 57%

Improving Multimodal LLMs Ability In Geometry Problem Solving, Reasoning, And Multistep Scoring

Avinash Anand, Raj Jaiswal, Abhishek Dharmadhikari, Atharva Marathe, Harsh Parimal Popat, Harshil Mital, Kritarth Prasad, Rajiv Ratn Shah, Roger Zimmermann

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16592 2024-10-23 cs.LG cs.CL cs.CY 57%

ViMGuard: A Novel Multi-Modal System for Video Misinformation Guarding

Andrew Kan, Christopher Kan, Zaid Nabulsi

专题命中 多模态RAG :retrieval augmented generation(abstract);分类 cs.CL

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13510 2024-10-18 cs.CL cs.CV 57%

GeoCoder: Solving Geometry Problems by Generating Modular Code through Vision-Language Models

Aditya Sharma, Aman Dalmia, Mehran Kazemi, Amal Zouaq, Christopher J. Pal

专题命中 多模态RAG :RAG(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18202 2024-10-01 cs.AI cs.MM 57%

WorldGPT: Empowering LLM as Multimodal World Model

Zhiqi Ge, Hongzhe Huang, Mingze Zhou, Juncheng Li, Guoming Wang, Siliang Tang, Yueting Zhuang

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.AI

Comments update v2

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11281 2024-09-18 cs.IR 57%

Beyond Relevance: Improving User Engagement by Personalization for Short-Video Search

Wentian Bao, Hu Liu, Kai Zheng, Chao Zhang, Shunyu Zhang, Enyun Yu, Wenwu Ou, Yang Song

专题命中 多模态RAG :dense retrieval(abstract);分类 cs.IR

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06450 2024-09-11 cs.RO cs.AI cs.ET 57%

Multimodal Large Language Model Driven Scenario Testing for Autonomous Vehicles

Qiujing Lu, Xuanhan Wang, Yiwei Jiang, Guangming Zhao, Mingyue Ma, Shuo Feng

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14723 2024-08-28 cs.CV cs.IR 57%

Snap and Diagnose: An Advanced Multimodal Retrieval System for Identifying Plant Diseases in the Wild

Tianqi Wei, Zhi Chen, Xin Yu

专题命中 多模态RAG :retriever(abstract);分类 cs.IR

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09272 2024-07-26 cs.CV cs.AI cs.SD eess.AS 57%

Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos

Changan Chen, Puyuan Peng, Ami Baid, Zihui Xue, Wei-Ning Hsu, David Harwath, Kristen Grauman

专题命中 多模态RAG :retrieval-augmented generation(abstract);分类 cs.AI

Comments Project page: https://vision.cs.utexas.edu/projects/action2sound. ECCV 2024 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15218 2024-02-26 cs.CR cs.CL cs.CV 57%

BSPA: Exploring Black-box Stealthy Prompt Attacks against Image Generators

Yu Tian, Xiao Yang, Yinpeng Dong, Heming Yang, Hang Su, Jun Zhu

专题命中 多模态RAG :retriever(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19787 2024-01-08 cs.CV cs.AI 57%

DeepMerge: Deep-Learning-Based Region-Merging for Image Segmentation

Xianwei Lv, Claudio Persello, Wangbin Li, Xiao Huang, Dongping Ming, Alfred Stein

专题命中 多模态RAG :RAG(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13820 2023-08-29 cs.IR 57%

Video and Audio are Images: A Cross-Modal Mixer for Original Data on Video-Audio Retrieval

Zichen Yuan, Qi Shen, Bingyi Zheng, Yuting Liu, Linying Jiang, Guibing Guo

专题命中 多模态RAG :retriever(abstract);分类 cs.IR

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09699 2023-08-21 cs.CV cs.CL 57%

PromptCap: Prompt-Guided Task-Aware Image Captioning

Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi, Noah A Smith, Jiebo Luo

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.CL

Comments Accepted to ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03512 2023-05-08 cs.CL 57%

Building Multimodal AI Chatbots

Min Young Lee

专题命中 多模态RAG :retriever(abstract);分类 cs.CL

Comments Bachelor's thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11507 2023-01-30 cs.CV cs.CL cs.LG 57%

Semi-Parametric Video-Grounded Text Generation

Sungdong Kim, Jin-Hwa Kim, Jiyoung Lee, Minjoon Seo

专题命中 多模态RAG :retriever(abstract);分类 cs.CL

Comments Preprint (16 pages, 5 figures)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08554 2022-10-18 cs.CV cs.CL 57%

COFAR: Commonsense and Factual Reasoning in Image Search

Prajwal Gatti, Abhirama Subramanyam Penamakuri, Revant Teotia, Anand Mishra, Shubhashis Sengupta, Roshni Ramnani

专题命中 多模态RAG :knowledge retrieval(abstract);分类 cs.CL

Comments Accepted in AACL-IJCNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.10858 2022-08-24 cs.IR 57%

VILT: Video Instructions Linking for Complex Tasks

Sophie Fischer, Carlos Gemmell, Iain Mackie, Jeffrey Dalton

专题命中 多模态RAG :dense retrieval(abstract);分类 cs.IR

Comments 7 pages, IMuR Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.11912 2021-08-27 cs.CV cs.CL cs.MM 57%

Similar Scenes arouse Similar Emotions: Parallel Data Augmentation for Stylized Image Captioning

Guodun Li, Yuchen Zhai, Zehao Lin, Yin Zhang

专题命中 多模态RAG :retriever(abstract);分类 cs.CL

Comments Accepted at ACM Multimedia (ACMMM) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.00265 2021-05-24 cs.CV cs.CL cs.LG 57%

VisualSparta: An Embarrassingly Simple Approach to Large-scale Text-to-Image Search with Weighted Bag-of-words

Xiaopeng Lu, Tiancheng Zhao, Kyusong Lee

专题命中 多模态RAG :vector search(abstract);分类 cs.CL

Comments Accepted to ACL2021 (10 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏