RKSC: Reasoning-Aware KV Cache Sharing and Confident Early Exit for Multi-Step LLM Inference
RKSC:面向多步LLM推理的感知推理的KV缓存共享与自信提前退出
机构 * Anirudh Sekar(安尼鲁德·塞卡)
专题命中 长上下文与记忆 :LLM(title,title_cn);分类 cs.CL、cs.AI、cs.LG
AI总结 提出RKSC框架,通过注意力相似性KV共享、置信门控提前退出和推理选择性块缓存管理,消除多分支LLM推理中的结构冗余,实现平均3.008倍加速,错误率仅0.37%。
Comments Accepted to the ICML 2026 Workshop on Statistical Frameworks for Uncertainty in Agentic Systems