arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

RAG / 检索增强生成

检索增强生成、向量检索、知识库问答和面向大模型的搜索系统。

共收录 8587 信号源:cs.IR, cs.CL, cs.AI, cs.DB

1. 多模态RAG 552 篇

2506.06962 2025-06-17 cs.CV 79%

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation

Jingyuan Qi, Zhiyang Xu, Qifan Wang, Lifu Huang

机构 * Virginia Tech(弗吉尼亚理工大学) Meta UC Davis(加州大学戴维斯分校)

专题命中 多模态RAG :RAG(title,abstract);retrieval augmented generation(comments)

Comments Image Generation, Retrieval Augmented Generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13706 2026-08-19 cs.CL cs.AI 版本更新 79%

CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA

CLAIR-Fin:用于跨模态金融问答中声明级验证与自适应辩论的对抗性多智能体框架

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Jubayer Al Mahmud, M. F. Mridha, Md. Alam Hossain

机构 * Ahsanullah University of Science and Technology(阿萨努拉科技大学) Jashore University of Science and Technology(杰索尔科技大学) American International University - Bangladesh(孟加拉国美国国际大学)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出CLAIR-Fin九智能体框架,针对跨模态金融问答的声明级验证与自适应辩论,在BB-FinQA-X数据集上提升了模型忠实度,且弃权比例合理,优于相关基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03836 2026-07-07 cs.CV cs.AI cs.CL 新提交 79%

When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts

何时越简单越好:评估中世纪拉丁文手稿的翻译管道

Nguyen Kim Hai Bui, Md. Easin Arafat, Tamás Gábor Orosz, Mufti Mahmud

机构 * Eötvös Loránd University(厄特沃什·罗兰大学) King Fahd University of Petroleum and Minerals(法赫德国王石油与矿产大学)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 针对历史手稿翻译难题,提出评估中世纪拉丁文手稿图像到翻译流程的框架。通过CATMuS数据集对比发现领域特定OCR模型优势,引入新数据集IPC,实验揭示简单管道表现更佳,为低资源历史场景翻译系统部署提供基准与指导。

Comments 17 pages, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28607 2026-05-28 cs.AI cs.CL 79%

Adaptive Multimodal Agents-Based Framework for Automatic Workflow Execution

基于自适应多智能体框架的自动工作流执行

Susanna Cifani, Mario Luca Bernardi, Marta Cimitile

机构 * Sapienza University of Rome(罗马萨皮恩扎大学) Department of Engineering University of Sannio(萨尼奥大学工程系) Faculty of Jurisprudence Unitelma Sapienza University(法理学院萨皮恩扎大学)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 提出一种多模态多智能体框架,通过离线构建拓扑知识库和在线自适应检索增强生成与闭环协作验证,实现自动工作流执行。

Comments Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses. Accepted for publication at the 2026 IEEE International Conference on Evolving and Adaptive Intelligent Systems (EAIS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18774 2026-05-20 cs.IR cs.AI 79%

M3DocDep: Multi-modal, Multi-page, Multi-document Dependency Chunking with Large Vision-Language Models

M3DocDep: 多模态、多页、多文档依赖分块方法基于大视觉-语言模型

Joongmin Shin, Jeongbae Park, Jaehyung Seo, Heuiseok Lim

机构 * Human-inspired AI Research, Korea University(韩国大学人智AI研究所) Computer Science and Engineering, Konkuk University(konkuk大学计算机科学与工程系) Department of Computer Science and Engineering, Korea University(韩国大学计算机科学与工程系)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 本文提出M3DocDep,一种基于大视觉-语言模型的多模态、多页、多文档依赖分块方法,通过恢复块级依赖并构建分块,提高了长多页多模态文档的检索和问答质量。

Comments Accepted to CVPR2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16275 2026-05-19 cs.CY cs.AI cs.CL cs.MM 79%

AI Slop or AI-enhancement? Student perceptions of AI-generated media for an English for Academic Purposes course

AI 产出物还是AI增强?英语学术用途课程中学生对AI生成媒体的看法

David James Woo, Deliang Wang, Kai Guo

机构 * Everwrite Limited(Everwrite有限公司) Faculty of Education, The University of Hong Kong(香港大学教育学院) Faculty of Education, The Chinese University of Hong Kong(香港中文大学教育学院)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨了AI生成内容在EAP课程中的教学效果,通过混合方法分析发现学生偏好视觉化内容,视频与学业表现正相关,但高认知负荷与成绩负相关,表明需合理设计内容以提升学习效果。

Comments 23 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11864 2026-05-13 cs.IR cs.AI cs.CV cs.MM 79%

Very Efficient Listwise Multimodal Reranking for Long Documents

非常高效的长文档多模态重排序方法

Yiqun Sun, Pengfei Wei, Lawrence B. Hsieh

机构 * Magellan Technology Research Institute (MTRI)(马杰拉技术研究院(MTRI))

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR、cs.AI

AI总结 本文提出ZipRerank,通过轻量级查询-图像早期交互机制和单次前向传递消除自回归解码,实现高效多模态重排序,实验表明其在MMDocIR基准上性能优异且显著降低LLM推理延迟。

Comments To appear in ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07241 2026-07-10 cs.SD 新提交 78%

Rag Classification of Tagore Songs using Symbolic Music Notation and Novel Weighted Distance Measures

使用符号音乐记谱法和新型加权距离度量对泰戈尔歌曲进行拉格分类

Chandan Misra, Swarup Chattopadhyay

机构 * XIM University(西姆大学)

专题命中 多模态RAG :RAG(title,abstract)

AI总结 该研究针对罗宾德拉·桑吉特歌曲拉格识别难题,将其转化为监督分类问题,利用符号乐谱记谱法构建数据集,探讨多种距离度量,引入加权欧几里得距离,在k近邻框架下改进拉格分类,更好捕捉旋律特征。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20414 2026-05-26 eess.AS 78%

PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding

PlanRAG-Audio:面向长音频理解的规划与检索增强生成

Masao Someki, Chien-yu Huang, Siddhant Arora, Samuele Cornell, Markus Müller, Nathan Susanj, Rupak V Swaminathan, Grant P Strimel, Jing Liu, Shinji Watanabe

专题命中 多模态RAG :retrieval augmented generation(title);retrieval-augmented generation(abstract)

AI总结 提出PlanRAG-Audio框架,通过规划查询所需模态和时间跨度并仅检索相关信息,实现长音频的高效推理,提升准确率并稳定性能。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04372 2026-04-07 cs.CV 78%

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning

图到帧RAG:面向无训练和可审计的视频推理的视觉空间知识融合

Songyuan Yang, Weijiang Yu, Ziyu Liu, Guijian Tang, Wenjing Yang, Huibin Tan, Nong Xiao

机构 * National University of Defense Technology(国防科技大学) Sun Yat-sen University(中山大学)

专题命中 多模态RAG :RAG(title,abstract)

AI总结 本文提出G2F-RAG,通过视觉空间知识融合提升视频推理的可解释性和效率,减少认知负担并保留可追溯的证据轨迹。

Comments Accepted at CVPR 2026. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02258 2026-03-24 cs.CV 78%

Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning

病理代理RAG:通过强化学习实现多模态代理检索增强生成用于病理学视觉语言模型

Wenchuan Zhang, Jingru Guo, Hengzhe Zhang, Penghao Zhang, Jie Chen, Shuwan Zhang, Zhang Zhang, Yuhao Yi, Hong Bu

专题命中 多模态RAG :retrieval-augmented generation(title);RAG(abstract)

AI总结 本文提出Patho-AgenticRAG,通过强化学习实现多模态代理检索增强生成,解决病理学视觉语言模型在高分辨率、复杂组织结构和临床语义上的挑战,提升诊断准确性。

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(35): 29921-29929, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23483 2025-12-30 cs.CV 78%

TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding

TV-RAG:一种具有时间意识和语义熵权的长视频检索与理解框架

Zongsheng Cao, Yangfan He, Anran Liu, Feng Chen, Zepeng Wang, Jun Xie

机构 * Researcher(研究者)

专题命中 多模态RAG :RAG(title,abstract)

AI总结 TV-RAG通过时间衰减检索和熵加权关键帧采样,提升长视频检索与理解性能,无需重新训练即可集成至现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05872 2025-12-10 cs.CV 78%

Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection

领域-RAG:跨领域少样本目标检测的检索引导组合图像生成

Yu Li, Xingyu Qiu, Yuqian Fu, Jie Chen, Tianwen Qian, Xu Zheng, Danda Pani Paudel, Yanwei Fu, Xuanjing Huang, Luc Van Gool, Yu-Gang Jiang

机构 * Fudan University(复旦大学) Fuzhou University(福州大学) East China Normal University(华东师范大学) HKUST(GZ)(香港科技大学(广州))

专题命中 多模态RAG :RAG(title,abstract)

AI总结 Domain-RAG通过检索引导的组合图像生成方法,解决跨领域少样本目标检测中的领域对齐与背景生成问题,实现高质量样本生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07538 2025-09-11 cs.CV 78%

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

Peijin Xie, Shun Qian, Bingquan Liu, Dexin Wang, Lin Sun, Xiangzheng Zhang

机构 * IEEE

专题命中 多模态RAG :RAG(title,abstract)

Comments 5 pages, 4 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00751 2025-09-03 cs.CV 78%

EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions

Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science, VNU-HCM(越南国家大学科学学院)

专题命中 多模态RAG :retriever(title,abstract)

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06496 2025-08-12 cs.CV cs.MA 78%

Med-GRIM: Enhanced Zero-Shot Medical VQA using prompt-embedded Multimodal Graph RAG

Rakesh Raj Madavan, Akshat Kaimal, Hashim Faisal, Chandrakala S

机构 * Shiv Nadar University Chennai(施瓦斯纳大学钦奈)

专题命中 多模态RAG :RAG(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08576 2025-03-12 cs.CV 78%

RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding

Xichen Tan, Yunfan Ye, Yuanjing Luo, Qian Wan, Fang Liu, Zhiping Cai

专题命中 多模态RAG :RAG(title,abstract)

Comments 37 pages, 36 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03340 2024-08-08 cs.MM cs.CV 78%

An Empirical Comparison of Video Frame Sampling Methods for Multi-Modal RAG Retrieval

Mahesh Kandhare, Thibault Gisselbrecht

专题命中 多模态RAG :RAG(title,abstract)

Comments 19 pages, 24 figures (65 images)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18386 2026-08-20 cs.CV cs.AI 新提交 77%

TTSD-FAR: Test-Time Self-Distillation with Fisher-Anchored Restoration for Missing-Modality Emotion Recognition in LVLMs

TTSD-FAR:用于大视频语言模型中缺失模态情感识别的带Fisher锚定恢复的测试时自蒸馏

Muhammad Haseeb Aslam, Alessandro Koerich, Marco Pedersoli, Ali Etemad, Eric Granger

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval augmented generation(abstract);分类 cs.AI

AI总结 针对大视频语言模型测试时的模态缺失问题,提出带Fisher锚定恢复的测试时自蒸馏框架,在多数据集模态缺失场景下性能优于基线方法且保持稳定。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.25041 2026-07-29 cs.IR cs.CR cs.CV cs.LG 新提交 77%

ScoreShield: Differentially Private Release of Similarity Scores

ScoreShield:相似度分数的差分隐私发布

Behrooz Razeghi, Parsa Rahimi

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR

AI总结 研究生物识别等应用中相似度分数隐私保护问题,提出ScoreShield先扰动后投影机制,添加校准高斯噪声并投影到可行性集,满足(ε,δ)-DP,给出效用保证,在多任务中评估,改进了风险的n依赖性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01814 2026-07-03 cs.AI 新提交 77%

MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support

MMIR-TCM:面向中医临床决策支持的记忆集成多模态推理与检索

Lihui Luo, Joongwon Chae, Ziyan Chen, Yang Liu, Siyi Cheng, Weihan Gao, Zelin Zeng, Xiaoming Yin, Samaneh Beheshti Kashi, Dongmei Yu, Lian Zhang, Jing Sui, Zeming Liang, Jiansong Ji, Peter E. Lobie, Peiwu Qin

机构 * Institute of Biopharmaceutics and Health Engineering, Tsinghua Shenzhen International Graduate School(清华深圳国际研究生院生物医药与健康工程研究院) Chinese Medicine Guangdong Laboratory(广东省中医药实验室) Beijing Normal University(北京师范大学) First Hospital of Hebei Medical University(河北医科大学第一医院) Lishui Hospital of Zhejiang University(浙江大学丽水医院) XiaoMing TCM Hospital(小明中医医院)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 提出MMIR-TCM框架,结合多模态大模型、记忆增强分割与检索增强生成,解决中医舌诊中的主观性和语义鸿沟问题,在MedTCM数据集上优于GPT-4o等模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04080 2026-06-30 cs.IR 版本更新 77%

Caption Injection for Optimization in Generative Search Engine

生成搜索引擎中的标题注入用于优化

Xiaolu Chen, Jie Bao, Haojie Wu, Zhen Chen, Yong Liao

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.IR

AI总结 本文提出Caption Injection,一种多模态G-SEO方法,通过提取图像标题并注入文本内容,提升生成搜索中的主观可见性,实验表明其在G-EVAL指标下优于文本-only基线。

Comments 24 pages, 4 figures, ECML PKDD 2026 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27974 2026-06-29 cs.CV cs.AI 新提交 77%

ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering

ProMSA: 渐进式多模态搜索智能体用于基于知识的视觉问答

ZhengXian Wu, Hangrui Xu, Kai Shi, Zhuohong Chen, Yunyao Yu, Chuanrui Zhang, Zirui Liao, Jun Yang, Zhenyu Yang, Haonan Lu, Haoqian Wang

机构 * OPPO AI Center, OPPO Inc. China(OPPO AI中心,OPPO公司) The Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Nanyang Technological University, Singapore(新加坡南洋理工大学)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retriever(abstract);分类 cs.AI

AI总结 提出渐进式多模态搜索智能体ProMSA,通过迭代选择图像搜索、文本搜索或停止,在显式工具调用预算和去重机制下,结合序列级强化学习优化,提升基于知识的视觉问答的检索和端到端准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18385 2026-06-18 cs.AI 新提交 77%

CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework

CaVe-VLM-CoT:一种可解释的视觉-语言模型框架

Sneha Rao, Shaina Raza, Dhanesh Ramachandram

机构 * Vector Institute(向量研究所)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retriever(abstract);分类 cs.AI

AI总结 提出CaVe-VLM-CoT框架,通过五阶段闭环流水线(提取器、检索器、求解器、引用注入器、验证器)实现证据推理,并引入CaVeScore复合指标评估检索质量、引用忠实度和跨模态基础,在ScienceQA和MMMU上取得性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08976 2026-06-16 cs.CV cs.DC cs.IR 版本更新 77%

MIRAGE: Runtime Scheduling for Multi-Vector Image Retrieval with Hierarchical Decomposition

MIRAGE:基于层次分解的多向量图像检索运行时调度

Maoliang Li, Ke Li, Yaoyang Liu, Jiayu Chen, Zihao Zheng, Yinjun Wu, Chenchen Liu, Xiang Chen

机构 * School of Computer Science, Peking University(北京大学计算机科学学院) School of Electronics Engineering and Computer Science, Peking University(北京大学电子工程与计算机科学学院) School of Information, Renmin University of China(中国人民大学信息学院) School of Integrated Circuit Science and Engineering, Beihang University(北京航空航天大学集成电路科学与工程学院)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval augmented generation(abstract);分类 cs.IR

AI总结 提出MIRAGE框架,通过层次化分解和跨层次相似性一致性减少冗余计算,实现多向量图像检索的精度提升和3.5倍计算加速。

Comments Will appear in DAC'2026, camera ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14113 2026-05-29 cs.CV cs.AI cs.LG cs.MA 77%

ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows

ProtoMedAgent: 通过隐私感知的智能体工作流实现多模态临床可解释性

Alvaro Lopez Pellicer, Plamen Angelov, Marwan Bukhari, Yi Li, Eduardo Soares, Jemma Kerns

机构 * School of Computing and Communications(计算与通信学校) Lancaster University(兰卡斯特大学) Lancaster Medical School(兰卡斯特医学院) PUC-Rio(里约热内卢联邦大学) Puc-Behring Institute for AI(人工智能皮克林研究所)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 提出ProtoMedAgent框架,通过神经符号瓶颈和反射性Scribe-Critic循环约束生成过程,解决原型网络在临床报告中的语义结构缺失和检索谄媚问题,并引入k-匿名和ℓ-多样性隐私门控。

Comments CVR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04565 2026-05-11 cs.MA cs.CL 77%

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems

从独立大语言模型到整合智能:复合AI系统综述

Jiayi Chen, Junyi Ye, Guiling Wang

机构 * New Jersey Institute of Technology(新泽西理工学院)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.CL

AI总结 本文综述了复合AI系统,探讨了其整合大语言模型与外部组件的方法,分析了四种基础范式,并指出规模化、互操作性等挑战及未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03454 2026-05-11 cs.CV cs.AI 77%

Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles

在驾驶前思考:基于世界模型的多模态接地用于自动驾驶车辆

Haicheng Liao, Huanming Shen, Bonan Wang, Yongkang Li, Yihong Tang, Chengyue Wang, Dingyi Zhuang, Kehua Chen, Hai Yang, Chengzhong Xu, Zhenning Li

机构 * University of Macau(澳门大学) UESTC(电子科技大学) Purdue University(普渡大学) McGill University(麦吉尔大学) Massachusetts Institute of Technology(麻省理工学院) University of Washington(华盛顿大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本文提出ThinkDeeper框架,通过预测未来空间状态提升自动驾驶车辆的自然语言指令理解能力,结合超图引导解码器融合多模态输入,提出DrivePilot数据集并在多个基准测试中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16345 2026-05-04 cs.HC cs.AI 77%

Bridging the Experimental Last Mile: Digitizing Laboratory Know-How for Safe AI-Assisted Support

弥合实验最后一公里:数字化实验室知识以实现安全的人工智能辅助支持

Akira Miura, Yuki Sasahara, Momoka Demura, Yuji Masubuchi, Tetsuya Asai, Chikahiko Mitsui

机构 * Division of Applied Chemistry, Faculty of Engineering, Hokkaido University(北海道大学工学部应用化学科) Graduate School of Chemical Sciences and Engineering, Hokkaido University(北海道大学研究生院化学科学和工程学系) Graduate School of Information Science and Technology, Hokkaido University(北海道大学研究生院信息科学和技术学系) Quantum Nexus, Inc.

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本文提出一种结合视频、多模态AI和检索增强生成的AI助手,通过提取实验室特定知识来提升实验安全性,经评估显示其在指导和安全方面具有实用性。

Comments 32 pages in total (main 13 pages, appendix 19 pages), 2 main figures, 1 main table

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17765 2026-05-01 q-bio.QM cs.AI cs.CV 77%

Grounded Multimodal Retrieval-Augmented Drafting of Radiology Impressions Using Case-Based Similarity Search

基于案例相似性搜索的医学影像印象 grounded 多模态检索增强草稿生成

Himadri S Samanta

机构 * Independent AI Researcher(独立AI研究员)

专题命中 多模态RAG :RAG(abstract,abstract_cn);retrieval-augmented generation(abstract);分类 cs.AI

AI总结 本文提出一种多模态检索增强生成系统,结合对比图像-文本嵌入、基于案例的相似性检索和引用约束草稿生成,以生成具有临床依据的医学影像印象。

Comments 15 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏