发表机构
Universitat Rovira i Virgili(罗维拉-威尔吉利大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LiteRAG通过查询条件算法探索和推理链构建,在保证多跳问答质量的同时,大幅降低延迟、成本和令牌使用。
AI 中文摘要
基于图的检索可以改进多跳问答,但现有方法往往产生高昂的查询时成本,并生成分散、过大的上下文,从而降低生成效率。我们提出LiteRAG,一种基于图的检索方法,用查询条件下的算法探索和推理链上下文构建取代昂贵的检索时LLM控制。在DistComp(一个针对分布式系统论文的多跳检索基准)上,LiteRAG在评估方法中取得最高整体质量(0.798),同时相对于GraphRAG Global和DRIFT,每次查询延迟降低超过100倍,成本降低超过99%。在UltraDomain上,LiteRAG在整体质量上与LinearRAG相当,同时使用的令牌数减少约14倍。消融研究表明,LiteRAG的查询自适应阈值和社区感知枢纽惩罚是其令牌效率提升的主要驱动因素。
英文摘要
Graph-based retrieval can improve multi-hop question answering, but existing approaches often incur high query-time costs and produce diffuse, oversized contexts that reduce generation efficiency. We present LiteRAG, a graph-based retrieval method that replaces expensive retrieval-time LLM control with query-conditioned algorithmic exploration and reasoning-chain context construction. On DistComp, a benchmark for multi-hop retrieval over distributed-systems papers, LiteRAG attains the highest overall quality among the evaluated methods (0.798) while reducing per-query latency by over 100$\times$ and cost by over 99% relative to GraphRAG Global and DRIFT. On UltraDomain, it matches LinearRAG on overall quality while using about 14$\times$ fewer tokens. An ablation study indicates that LiteRAG's query-adaptive thresholding and community-aware hub penalization are the main drivers of its token-efficiency gains.
Comments16 pages, 2 figures