基于SSD稀疏KV存储的高效智能体LLM服务
Efficient Agentic LLM Serving over SSD-based Sparse KV Storage
浏览论文内容
中文总结 AI 辅助
针对智能体LLM服务中SSD稀疏KV存储的读取延迟问题,提出Janus框架,通过预测KV需求将读取移出关键路径并优化SSD效率,显著降低首词延迟。
中文摘要 AI 辅助
由大型语言模型(LLM)驱动的智能体会话通常在模型推理和工具使用之间交替进行,并在连续轮次中累积长期历史记录。高效地服务这些会话需要减少注意力计算并保留历史键值(KV)缓存以避免重新计算。近期,前沿开源LLM采用稀疏注意力机制,通过仅选择部分历史记录来减少计算量,而SSD为存储KV缓存提供了比CPU DRAM更廉价的替代方案。然而,稀疏KV选择依赖于模型推理过程中的临时中间值,这迫使SSD读取处于推理关键路径上。这些读取还因SSD中的碎片化访问和读写干扰而进一步变慢。为解决这些挑战,我们提出了Janus,一个面向稀疏注意力LLM的、以SSD为中心的KV存储的智能体服务框架。Janus专注于追加预填充(append prefill),该过程处理每轮新添加的输入,并承担大部分历史KV加载。为了将SSD读取移出关键路径,Janus在较早的中间值上运行模型自身的KV选择模块,无需额外训练即可预测KV需求。预测的读取与模型计算重叠,任何预测未命中都会在注意力执行前被获取以保持模型输出。为提高SSD效率,Janus合并相邻读取,将分散的KV页打包为CPU上的顺序写入,并在读取活跃时限制后台写入。在三个模型和三个智能体轨迹上,Janus在首词延迟(time to first token)方面比现有工作提升了1.57-3.69倍(平均1.22-1.85倍),同时保持解码效率。
英文摘要
Agentic sessions driven by Large language models (LLMs) often alternate between model inference and tool use, accumulating long histories across successive rounds. Serving these sessions efficiently requires reducing attention computation and retaining history key-value (KV) caches to avoid recomputation. Recently, frontier open-source LLMs adopt sparse attention to reduce computation by selecting only part of the history, while SSDs provide a cheaper alternative to CPU DRAM for storing KV caches. However, sparse KV selection depends on the ad hoc intermediate values during model inference, so it forces SSD reads to lie on the inference critical path. These reads are further slowed by fragmented accesses and read-write interference in SSDs. To address these challenges, we present Janus, an agentic serving framework for sparse attention LLMs with SSD-centric KV storage. Janus focuses on append prefill, which processes each round's newly added inputs and accounts for most history KV loading. To move SSD reads out of the critical path, Janus runs the model's own KV selection module on earlier intermediate values, predicting KV demand without additional training. The predicted reads overlap with model computation, and any prediction misses are fetched before attention executes to preserve model outputs. To improve SSD efficiency, Janus coalesces adjacent reads, packs scattered KV pages into sequential writes on the CPU, and limits background writes while reads are active. Across three models and three agentic traces, Janus outperforms existing works by up to 1.57-3.69 times (1.22-1.85 times on average) in terms of the time to first token latency, while maintaining decode efficiency.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Shanghai Innovation Institute(上海创新研究院)
- Shanghai Qiji Zhifeng Co Ltd(上海奇积智峰有限公司)
机构由 AI 辅助整理,请以论文原文为准。