arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越静态RAG:一种面向商品GPU高效长上下文推理的自适应三指标路由框架

Beyond Static RAG: An Adaptive, Tri-Metric Routing Framework for Efficient Long-Context Inference on Commodity GPUs

Saipraveen Vabbilisetty, Ajay Kumar Boddepalli, Deep Narayan Mishra, Shashank Kapadia, Haoan Wang, Anupriya Sharma

arXiv 2609.17564首次发表:更新:

AI 中文总结

针对商品GPU上RAG长上下文推理的压缩悖论,提出一种基于显存余量和延迟交叉点的三指标路由框架,实现0% OOM故障并提升5.2点F1。

AI 中文摘要

在NVIDIA T4(16 GB显存)等商品GPU上部署检索增强生成(RAG)时,会暴露一种我们称之为“压缩悖论”的实际失效模式:神经提示压缩可能增加键值(KV)缓存争用和预处理延迟,其代价超过生成时间的节省,而跳过压缩则可能导致长上下文场景下的内存不足(OOM)故障。我们识别了在内存预算紧张时,由vLLM服务的LLM与基于PyTorch的压缩器共同部署时出现的两种不同失效机制,并引入了三指标路由器(Tri-Metric Router),这是一种确定性的、无需训练的策略,可在原始(Raw)、神经(LLMLingua-2)和词法(BM25)流水线之间进行选择。该路由器使用三个CPU侧信号:空间复杂度($L$)、句法密度($\ ho_{key}$)和类型-标记比(TTR)。与先前仅基于语义的适配不同,我们的调度信号是硬件物理层面的,基于显存余量和延迟交叉点。阈值通过LongBench qasper上的性能分析校准,在T4上产生约4,332词的运行交叉点;我们的贡献在于这种校准方法,而非硬件特定的常数。在分布外保留集上,该方法实现了0%的OOM故障、88.5±4.4%的预言机对齐率以及49.3%的Combined F1,相比始终开启的词法压缩提升了5.2个百分点,且无需额外显存或训练成本。

英文摘要

Deploying retrieval-augmented generation (RAG) on commodity GPUs such as the NVIDIA T4 (16 GB VRAM) exposes a practical failure mode we call the Compression Paradox: neural prompt compression can add key-value (KV) cache contention and preprocessing latency that outweigh generation-time savings, while skipping compression can cause out-of-memory (OOM) failures on long contexts. We identify two distinct failure mechanisms when a vLLM-served LLM and a PyTorch-based compressor are co-deployed under tight memory budgets, and introduce the Tri-Metric Router, a deterministic, training-free policy that selects among Raw, Neural (LLMLingua-2), and Lexical (BM25) pipelines. The router uses three CPU-side signals: spatial complexity ($L$), syntactic density ($ρ_{key}$), and type-token ratio (TTR). Unlike prior semantic-only adaptation, our dispatch signal is hardware-physical, based on VRAM headroom and a latency crossover point. Thresholds are calibrated from profiling on LongBench qasper, yielding an operating crossover near 4,332 words on T4; our contribution is this calibration methodology rather than a hardware-specific constant. On out-of-distribution holdouts, the method achieves 0% OOM failures, 88.5 $\pm$ 4.4% oracle alignment, and 49.3% Combined F1, improving over always-on lexical compression by 5.2 points without additional VRAM or training cost.

CommentsThis Paper is accepted and presented at ICML Scale Workshop 2026. https://scale-icml-2026.github.io/accepted_papers.html (Paper ID :72)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑