arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 8633 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 557 篇

2511.06752 2025-11-11 cs.CV 50%

Med-SORA: Symptom to Organ Reasoning in Abdomen CT Images

You-Kyoung Na, Yeong-Jun Cho

机构 * Chonnam National University(全南国立大学)

专题命中 多模态RAG :RAG(abstract)

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05199 2025-11-10 cs.RO 50%

Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation

Yichen Zhu, Feifei Feng

机构 * Midea Group, AI Research Center(美的集团人工智能研究中心)

专题命中 多模态RAG :retriever(abstract)

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26386 2025-10-29 cs.CV 50%

PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer

Zhiwei Yang, Chen Gao, Mike Zheng Shou

机构 * Xidian University(西安电子科技大学) Show Lab, National University of Singapore(新加坡国立大学Show实验室)

专题命中 多模态RAG :RAG(abstract)

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16510 2025-10-21 q-bio.BM 50%

CryoDyna: Multiscale end-to-end modeling of cryo-EM macromolecule dynamics with physics-aware neural network

Chengwei Zhang, Shimian Li, Yihao Niu, Zhen Zhu, Sihao Yuan, Sirui Liu, Yi Qin Gao

专题命中 多模态RAG :RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19203 2025-09-24 cs.CV 50%

Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions

Ioanna Ntinou, Alexandros Xenos, Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

机构 * Queen Mary University of London(伦敦女王大学) Samsung AI Centre(三星人工智能中心) Technical University of Iași(伊阿苏技术大学)

专题命中 多模态RAG :retriever(abstract)

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01960 2025-09-23 cs.LG 50%

MPIC: Position-Independent Multimodal Context Caching System for Efficient MLLM Serving

Shiju Zhao, Junhao Hu, Rongxiao Huang, Jiaqi Zheng, Guihai Chen

机构 * State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) Nanjing University(南京大学) School of Computer Science(计算机学院)

专题命中 多模态RAG :retrieval-augmented generation(abstract)

Comments 17 pages, 13 figures, the second version

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13899 2025-09-18 cs.HC 50%

AI as a teaching tool and learning partner

Steven Watterson, Sarah Atkinson, Elaine Murray, Andrew McDowell

专题命中 多模态RAG :RAG(abstract)

Comments 6 Pages, 1 Figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01970 2025-08-29 cs.LG 50%

Improving Hospital Risk Prediction with Knowledge-Augmented Multimodal EHR Modeling

Rituparna Datta, Jiaming Cui, Zihan Guan, Vishal G. Reddy, Joshua C. Eby, Gregory Madden, Rupesh Silwal, Anil Vullikanti

机构 * Department of Computer Science, University of Virginia(大学计算机科学系) University of Virginia School of Medicine(弗吉尼亚大学医学院) Virginia Polytechnic Institute and State University(弗吉尼亚理工学院和州立大学) Biocomplexity Institute and Initiative, University of Virginia(大学生物复杂性研究所) Division of Infectious Diseases & International Health, University of Virginia School of Medicine(大学感染性疾病与国际卫生分会)

专题命中 多模态RAG :knowledge retrieval(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11974 2025-07-17 cs.RO 50%

A Review of Generative AI in Aquaculture: Foundations, Applications, and Future Directions for Smart and Sustainable Farming

Waseem Akram, Muhayy Ud Din, Lyes Saad Soud, Irfan Hussain

机构 * Khalifa University Center for Autonomous Robotic Systems (KUCARS), Khalifa University, United Arab Emirates(卡里法大学自主机器人系统中心(KUCARS)、卡里法大学、阿拉伯联合酋长国)

专题命中 多模态RAG :retrieval augmented generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02900 2025-07-11 cs.CV 50%

MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine

Yunfei Xie, Ce Zhou, Lang Gao, Juncheng Wu, Xianhang Li, Hong-Yu Zhou, Sheng Liu, Lei Xing, James Zou, Cihang Xie, Yuyin Zhou

机构 * Huazhong University of Science and Technology(华中科技大学) UC Santa Cruz(加州大学圣克ruz分校) Harvard University(哈佛大学) Stanford University(斯坦福大学)

专题命中 多模态RAG :retrieval-augmented generation(abstract)

Comments The dataset is publicly available at https://yunfeixie233.github.io/MedTrinity-25M/. Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18899 2025-06-24 cs.CV 50%

FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation

Kaiyi Huang, Yukun Huang, Xintao Wang, Zinan Lin, Xuefei Ning, Pengfei Wan, Di Zhang, Yu Wang, Xihui Liu

机构 * The University of Hong Kong(香港大学) Kuaishou Technology(快手科技) Microsoft Research(微软研究院) Tsinghua University(清华大学)

专题命中 多模态RAG :RAG(abstract)

Comments Project Page: https://filmaster-ai.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17837 2025-06-24 cs.CV 50%

Time-Contrastive Pretraining for In-Context Image and Video Segmentation

Assefa Wahd, Jacob Jaremko, Abhilash Hareendranathan

机构 * Department of Radiology(放射科部门)

专题命中 多模态RAG :retriever(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10533 2025-05-16 cs.CV cs.LG 50%

Enhancing Multi-Image Question Answering via Submodular Subset Selection

Aaryan Sharma, Shivansh Gupta, Samar Agarwal, Vishak Prasad C., Ganesh Ramakrishnan

机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔学院)

专题命中 多模态RAG :retriever(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05782 2025-03-18 cs.SD cs.CV cs.LG cs.MM eess.AS 50%

Sequential Contrastive Audio-Visual Learning

Ioannis Tsiamas, Santiago Pascual, Chunghsin Yeh, Joan Serrà

专题命中 多模态RAG :hybrid retrieval(abstract)

Comments ICASSP 2025. Version 1 contains more details

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02692 2025-03-05 cs.CE econ.GN q-fin.EC 50%

FinArena: A Human-Agent Collaboration Framework for Financial Market Analysis and Forecasting

Congluo Xu, Zhaobin Liu, Ziyang Li

专题命中 多模态RAG :RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21068 2025-03-03 cs.SE 50%

GUIDE: LLM-Driven GUI Generation Decomposition for Automated Prototyping

Kristian Kolthoff, Felix Kretzer, Christian Bartelt, Alexander Maedche, Simone Paolo Ponzetto

专题命中 多模态RAG :retrieval-augmented generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17038 2025-02-25 cs.MM 50%

Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction

Jiacheng Lu, Mingyuan Xiao, Weijian Wang, Yuxin Du, Zhengze Wu, Cheng Hua

专题命中 多模态RAG :retriever(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20725 2024-12-31 cs.CV 50%

Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling

Min Zhang, Zilin Wang, Liyan Chen, Kunhong Liu, Juncong Lin

专题命中 多模态RAG :retrieval-augmented generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09936 2024-12-16 cs.CV 50%

CaLoRAify: Calorie Estimation with Visual-Text Pairing and LoRA-Driven Visual Language Models

Dongyu Yao, Keling Yao, Junhong Zhou, Yinghao Zhang

专题命中 多模态RAG :RAG(abstract)

Comments Disclaimer: This work is part of a course project and reflects ongoing exploration in the field of vision-language models and calorie estimation. Findings and conclusions are subject to further validation and refinement

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00304 2024-11-04 cs.CV cs.MM 50%

Unified Generative and Discriminative Training for Multi-modal Large Language Models

Wei Chow, Juncheng Li, Qifan Yu, Kaihang Pan, Hao Fei, Zhiqi Ge, Shuai Yang, Siliang Tang, Hanwang Zhang, Qianru Sun

专题命中 多模态RAG :retrieval-augmented generation(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19720 2024-10-01 cs.CV 50%

FAST: A Dual-tier Few-Shot Learning Paradigm for Whole Slide Image Classification

Kexue Fu, Xiaoyuan Luo, Linhao Qu, Shuo Wang, Ying Xiong, Ilias Maglogiannis, Longxiang Gao, Manning Wang

专题命中 多模态RAG :knowledge retrieval(abstract)

Comments Accepted to NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04717 2024-05-09 cs.CV 50%

Remote Diffusion

Kunal Sunil Kasodekar

专题命中 多模态RAG :RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03427 2023-10-03 eess.IV cs.CV cs.LG 50%

Merging-Diverging Hybrid Transformer Networks for Survival Prediction in Head and Neck Cancer

Mingyuan Meng, Lei Bi, Michael Fulham, Dagan Feng, Jinman Kim

专题命中 多模态RAG :RAG(abstract)

Comments Early Accepted at International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI 2023)

Journal ref International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI), pp. 400-410, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09564 2023-08-21 cs.CV 50%

Deep Equilibrium Object Detection

Shuai Wang, Yao Teng, Limin Wang

专题命中 多模态RAG :RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16995 2022-12-01 eess.IV cs.CV 50%

A hybrid motion estimation technique for fisheye video sequences based on equisolid re-projection

Andrea Eichenseer, Michel Bätz, Jürgen Seiler, André Kaup

专题命中 多模态RAG :vector search(abstract)

Journal ref IEEE International Conference on Image Processing (ICIP), 2015, pp. 3565-3569

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.00775 2022-01-24 cs.CV 50%

PP-ShiTu: A Practical Lightweight Image Recognition System

Shengyu Wei, Ruoyu Guo, Cheng Cui, Bin Lu, Shuilong Dong, Tingquan Gao, Yuning Du, Ying Zhou, Xueying Lyu, Qiwen Liu, Xiaoguang Hu, Dianhai Yu, Yanjun Ma

专题命中 多模态RAG :vector search(abstract)

Comments 9 pages, 5 figures, 9 tables. arXiv admin note: text overlap with arXiv:2109.03144

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.08949 2021-12-17 cs.CV cs.LG 50%

Slot-VPS: Object-centric Representation Learning for Video Panoptic Segmentation

Yi Zhou, Hui Zhang, Hana Lee, Shuyang Sun, Pingjun Li, Yangguang Zhu, ByungIn Yoo, Xiaojuan Qi, Jae-Joon Han

专题命中 多模态RAG :retriever(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.00945 2021-02-23 cs.CV 50%

Two-Stage Single Image Reflection Removal with Reflection-Aware Guidance

Yu Li, Ming Liu, Yaling Yi, Qince Li, Dongwei Ren, Wangmeng Zuo

专题命中 多模态RAG :RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.11193 2020-12-22 cs.CV 50%

Image Translation via Fine-grained Knowledge Transfer

Xuanhong Chen, Ziang Liu, Ting Qiu, Bingbing Ni, Naiyuan Liu, Xiwei Hu, Yuhan Li

专题命中 多模态RAG :knowledge retrieval(abstract)

Comments Submitted to CVPR2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.05544 2020-11-17 cs.LG cs.CV stat.ML 50%

Superpixel Image Classification with Graph Attention Networks

Pedro H. C. Avelar, Anderson R. Tavares, Thiago L. T. da Silveira, Cláudio R. Jung, Luís C. Lamb

专题命中 多模态RAG :RAG(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏