面向三值大语言模型的带符号数位键值缓存的统一查表推理
Unified Lookup-Table Inference with Signed-Digit K/V Caches for Ternary LLMs
浏览论文内容
中文总结 AI 辅助
针对三值大语言模型注意力机制中K/V缓存与权重投影的精度不匹配问题,提出带符号数位K/V缓存的统一查表推理方法,经实验验证可提升缓存、质量与硬件效率。
中文摘要 AI 辅助
三值大语言模型(Ternary LLMs)使以权重为主的投影操作变得紧凑高效,但注意力机制仍存在不匹配问题:其键值(K/V)缓存是在线生成的,通常由单独的高精度引擎处理,仅压缩该缓存无法解决这种不匹配。为使用与三值投影相同的查表机制执行注意力,一次归约中累积的值必须保留兼容的表示形式和缩放比例;在因果解码期间,键和值的这一要求也存在差异,因为新生成的值可能属于未完成的缓存块。本研究提出一种面向三值大语言模型的统一查表推理方法,将运行时K/V状态存储为围绕注意力归约结构组织的缩放多平面带符号数位,生成的数位平面可直接由激活衍生的表使用,避免了缓存存储与注意力计算之间的密集K/V物化。该设计结合了在线K/V形成、不完整值块的有界处理,以及线性投影与注意力的共享多流数据通路;通过约束引导的搜索为目标质量效率权衡选择表示形式和执行策略。在原生及后训练三值模型上开展的实验,从缓存容量、模型质量和硬件效率等方面验证了该方法的有效性。
英文摘要
Ternary LLMs make their weight-dominated projections compact and efficient, but attention remains a mismatch: its K/V cache is created online and is typically processed by a separate higher-precision engine. Compressing this cache alone does not resolve the mismatch. To execute attention with the same lookup-table machinery as ternary projections, values accumulated in one reduction must retain a compatible representation and scale. This requirement also differs for keys and values during causal decoding, because newly generated values may belong to an unfinished cache block. This work develops a unified lookup-table inference approach for ternary LLMs. It stores runtime K/V states as scaled multi-plane signed digits organized around the reduction structure of attention. The resulting digit planes are consumed directly by activation-derived tables, avoiding dense K/V materialization between cache storage and attention computation. The design combines online K/V formation, bounded handling of incomplete value blocks, and a shared multi-stream datapath for Linear projections and attention. A constraint-guided search selects the representation and execution policy for a target quality--efficiency trade-off. Experiments on native and post-training ternary models validate the approach across cache capacity, model quality, and hardware efficiency.