Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving
Lethe:面向推理密集型大语言模型服务的层和时间自适应KV缓存剪枝
专题命中 效率与部署 :LLM(title);large language model(abstract);language model(abstract);分类 cs.LG
AI总结 Lethe通过引入时空自适应的KV缓存管理框架,在推理密集型大语言模型服务中实现高效的缓存剪枝,提升吞吐量的同时保持生成质量。
Comments aaai26 camera-ready version, 10 pages