P-RAG:结合LoRA和选择性CoT的增强型参数RAG
P-RAG: Prompt-Enhanced Parametric RAG with LoRA and Selective CoT for Biomedical and Multi-Hop QA
- University of Washington(华盛顿大学)
- Swarthmore College(斯沃斯莫尔学院)
- Sichuan Agricultural University(四川农业大学)
- Skyline High School(斯凯莱恩高中)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
P-RAG通过LoRA和选择性CoT提升生物医学问答的准确性和可扩展性。
AI中文摘要:
大型语言模型(LLMs)展现出显著的能力,但仍然受限于其对静态训练数据的依赖。检索增强生成(RAG)通过在推理过程中检索外部知识来解决这一限制,尽管它仍然严重依赖知识库的质量。为了探索潜在的改进,我们评估了三种RAG变体——标准RAG、DA-RAG以及我们提出的一种增强型参数RAG(P-RAG),这是一种将参数化知识嵌入LLM并检索证据的混合架构,由链式推理(CoT)提示和低秩适应(LoRA)微调引导。使用通过LoRA微调的LLaMA-3.2-1B-Instruct模型,在PubMedQA和2WikiMultiHopQA数据集上进行评估。P-RAG在PubMedQA上比标准RAG在F1分数上高出10.47个百分点(93.33% vs. 82.86%;12.64%相对)。在2WikiMultiHopQA上,P-RAG的总体得分几乎翻倍(33.44% vs. 17.83%),并在Compare子集上达到44.03%(其中Bridge为42.74%,Inference为21.84%,Compose为8.60%)。CoT提示显著提高了多跳推理,但对更简单的单跳查询效果混合。这些发现凸显了P-RAG在准确、可扩展和上下文适应的生物医学问答中的潜力。我们的贡献包括:(1)基于LoRA的LLaMA-3.2-1B-Instruct模型对生物医学问答的微调;(2)引入P-RAG及其链式推理提示;(3)在PubMedQA和2WikiMultiHopQA上取得最先进的结果。
英文摘要:
Large Language Models (LLMs) demonstrate remarkable capabilities but remain limited by their reliance on static training data. Retrieval-Augmented Generation (RAG) addresses this constraint by retrieving external knowledge during inference, though it still depends heavily on knowledge base quality. To explore potential improvements, we evaluated three RAG variants-Standard RAG, DA-RAG, and our proposed Prompt-Enhanced Parametric RAG (P-RAG), a hybrid architecture that integrates parametric knowledge within the LLM and retrieved evidence, guided by Chain-of-Thought (CoT) prompting and Low-Rank Adaptation (LoRA) fine-tuning-on both general and biomedical datasets. Using LLaMA-3.2-1B-Instruct fine-tuned via LoRA, we evaluate on PubMedQA and 2WikiMultihopQA. P-RAG outperforms Standard RAG on PubMedQA by 10.47 percentage points in F1 (93.33% vs. 82.86%; 12.64% relative). On 2WikiMultihopQA, P-RAG nearly doubles the overall score vs. Standard RAG (33.44% vs. 17.83%) and achieves 44.03% on the Compare subset (with 42.74% Bridge, 21.84% Inference, 8.60% Compose). CoT prompting substantially improves multi-hop reasoning but yields mixed results for simpler, single-hop queries. These findings underscore P-RAG's potential for accurate, scalable, and contextually adaptive biomedical question answering. Our contributions include: (1) LoRA-based fine-tuning of LLaMA-3.2-1B-Instruct for biomedical QA, (2) introduction of P-RAG with Chain-of-Thought prompting, and (3) state-of-the-art results on PubMedQA and 2WikiMultihopQA.