ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
ARKV:在有限内存预算下为LLMs长上下文推理实现自适应和资源高效的KV缓存管理
专题命中 数学推理 :reasoning(abstract);math reasoning(abstract);分类 cs.AI
AI总结 ARKV通过动态分配精度级别,实现LLMs长上下文推理中的高效内存管理,减少内存使用并保持高精度。
Comments Accepted in ACM/IEEE CCGRID 2025 conference