arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19535cs.AIcs.CLcs.DCcs.IRcs.PF

从检索上下文到运行时控制:基于边缘设备的检索增强生成(RAG)的自适应压缩

From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG

Zlatan Feric, Amir Taherin, Yanzhi Wang, David Kaeli

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对边缘RAG的上下文压缩问题,提出遥测驱动的自适应压缩方案,通过实验发现中等压缩可显著降低能耗且质量损失极小,主张基于工作负载和边缘遥测动态管理压缩。

中文摘要 AI 辅助

检索增强生成(RAG)通过将生成过程建立在外部段落基础上,改进了语言模型的响应,但会带来开销:检索到的上下文会延长提示词,增加预填充工作量、KV缓存占用量、内存流量、延迟和能耗。上下文压缩是一种自然的解决方案,可在生成前修剪检索到的文本。然而,最先进的上下文压缩方法通常使用固定的压缩预算,或在推理时应用离线选择的速率。这种静态视角忽略了工作负载变化和边缘设备的实时状态。在边缘片上系统(SoC)上,压缩并非免费:压缩器本身在同一SoC上运行,会消耗延迟和能耗,可能抵消生成带来的任何节省。本文基于实验证据,提出了边缘RAG中遥测驱动的自适应压缩的设想。我们使用Llama和Qwen生成器、Natural Questions和HotpotQA数据集以及LLMLingua-2压缩,在NVIDIA Jetson AGX Thor上表征了压缩的权衡关系。测量结果显示,对于较大的模型,生成过程占据了RAG开销的主要部分,在7B-8B生成器中,生成过程约占每个查询延迟的90%和GPU能耗的91%。探索压缩率的影响后,我们发现了一个自适应操作区域:轻度压缩可能会错过节能机会,过度激进的压缩会损害推理质量,而中等程度的压缩可将GPU能耗最多降低53.2%,SoC能耗最多降低48.2%,且质量损失可忽略不计。我们主张制定运行时策略,根据工作负载特征和边缘遥测动态管理压缩。

英文摘要

Retrieval-augmented generation (RAG) improves language-model responses by grounding generation in external passages, which comes with overhead: retrieved context lengthens the prompt, increasing prefill work, KV-cache footprint, memory traffic, latency, and energy. Context compression offers a natural remedy by pruning retrieved text before generation. However, state-of-the-art context-compression methods are typically used with a fixed compression budget, or with the rate selected offline and then applied at inference time. This static view ignores both workload variation and the live state of the edge device. On an edge SoC, compression is not free: the compressor itself runs on the same SoC and consumes latency and energy that can offset any generation savings. This paper proposes a vision for telemetry-informed adaptive compression in edge RAG, grounded in experimental evidence. We characterize the compression tradeoff on the NVIDIA Jetson AGX Thor using Llama and Qwen generators, Natural Questions and HotpotQA datasets, and LLMLingua-2 compression. Our measurements show that generation dominates the RAG budget for larger models, reaching roughly 90% of per-query latency and 91% of GPU energy for 7B-8B generators. Exploring the impact of the compression rate reveals an adaptive operating region: mild compression can miss energy opportunities, and overly aggressive compression can hurt inference quality. Intermediate compression can reduce GPU energy by up to 53.2%, and SoC energy by up to 48.2%, with negligible quality loss. We argue for runtime policies that dynamically manage compression, guided by workload features and edge telemetry.

发表机构

  • Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑