发表机构
State Key Laboratory of AI Safety; Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Meituan(人工智能安全国家重点实验室; 中国科学院计算技术研究所; 中国科学院大学; 美团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
R$^{2}$Adapter是一种轻量即插即用适配器,可动态分配查询至基础RAG与图RAG,减少图RAG使用量达59%且保持准确率,为混合RAG提供高效自适应方案
AI 中文摘要
检索增强生成(Retrieval-Augmented Generation, RAG)已成为利用非参数知识增强大型语言模型(Large Language Models, LLMs)的主流范式。基础RAG能高效处理简单查询,但在关系推理或多跳推理任务中表现不佳。基于图的RAG可缓解该问题,但会带来更高的推理复杂度和延迟。实际场景中,用户查询的复杂度差异显著,固定的RAG策略并非最优。然而,现有的文本-图混合RAG方法通常依赖启发式规则或基于LLM的路由,导致不必要的开销且过度依赖底层LLM。为解决这些挑战,我们提出R$^{2}$Adapter,这是一种轻量级的即插即用路由与重写适配器,旨在动态分配查询至基础RAG和基于图的RAG。通过仅路由那些真正能从基于图的推理中受益的查询,R$^{2}$Adapter减少了不必要的图检索开销。此外,对经图路由的不确定查询进行重写,以更好地凸显其多跳推理需求,在无需额外监督的情况下提升了检索质量。在三个多跳问答基准上开展的大量实验表明,R$^{2}$Adapter可将基于图的RAG的使用量降低多达59%,同时保持相当的答案准确率。该适配器与模型无关,可无缝集成到各类基础RAG和基于图的RAG流程中,为混合RAG系统提供了高效且自适应的解决方案。
英文摘要
Retrieval-Augmented Generation (RAG) has become a prevailing paradigm for enhancing Large Language Models (LLMs) with non-parametric knowledge. Vanilla RAG efficiently handles simple queries but struggles with relational or multi-hop reasoning. Graph-based RAG alleviates this issue but incurs higher inference complexity and latency. In practice, user queries can differ significantly in their complexity, rendering a fixed RAG strategy suboptimal. However, existing hybrid text-graph RAG methods typically rely on heuristic and LLM-based routing, resulting in unnecessary overhead and strong dependence on the underlying LLM. To address these challenges, we propose R$^{2}$Adapter, a lightweight plug-in Routing and Rewriting Adapter designed to allocate queries between vanilla and graph-based RAG dynamically. By routing only the queries that genuinely benefit from graph-based reasoning, R$^{2}$Adapter reduces unnecessary graph retrieval overhead. Additionally, uncertain graph-routed queries are rewritten to better expose their multi-hop reasoning requirements, improving retrieval quality without additional supervision. Extensive experiments on three multi-hop QA benchmarks demonstrate that R$^{2}$Adapter reduces graph-based RAG usage by up to 59% while maintaining comparable answer accuracy. This adapter is model-agnostic and can be seamlessly integrated into diverse vanilla and graph-based RAG pipelines, providing an efficient and adaptive solution for hybrid RAG systems.