Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
解开地址簿:通过稀疏自编码器解剖LLM键值缓存的稀疏语义结构
机构 * Beijing University of Posts(北京邮电大学) ; Baidu Inc.(百度公司)
专题命中 长上下文与记忆 :LLM(title);large language model(abstract);language model(abstract);分类 cs.LG
AI总结 本文提出STA-Attention框架,通过Top-K稀疏自编码器解剖LLM键值缓存的稀疏语义结构,实现可解释的语义原子分解,提升模型的可解释性和性能。