arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03229cs.AR

面向三值大语言模型的带符号数位键值缓存的统一查表推理

Unified Lookup-Table Inference with Signed-Digit K/V Caches for Ternary LLMs

Ziang Duan, Jiajun Wu, Zetian Chen, Hao Song, Yanwen Deng, Zixuan Shen, Nuobei Xie, Simo Wu, Bolun Wang, Peng Zhou, Chao Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对三值大语言模型注意力机制中K/V缓存与权重投影的精度不匹配问题,提出带符号数位K/V缓存的统一查表推理方法,经实验验证可提升缓存、质量与硬件效率。

中文摘要 AI 辅助

三值大语言模型(Ternary LLMs)使以权重为主的投影操作变得紧凑高效,但注意力机制仍存在不匹配问题:其键值(K/V)缓存是在线生成的,通常由单独的高精度引擎处理,仅压缩该缓存无法解决这种不匹配。为使用与三值投影相同的查表机制执行注意力,一次归约中累积的值必须保留兼容的表示形式和缩放比例;在因果解码期间,键和值的这一要求也存在差异,因为新生成的值可能属于未完成的缓存块。本研究提出一种面向三值大语言模型的统一查表推理方法,将运行时K/V状态存储为围绕注意力归约结构组织的缩放多平面带符号数位,生成的数位平面可直接由激活衍生的表使用,避免了缓存存储与注意力计算之间的密集K/V物化。该设计结合了在线K/V形成、不完整值块的有界处理,以及线性投影与注意力的共享多流数据通路;通过约束引导的搜索为目标质量效率权衡选择表示形式和执行策略。在原生及后训练三值模型上开展的实验,从缓存容量、模型质量和硬件效率等方面验证了该方法的有效性。

英文摘要

Ternary LLMs make their weight-dominated projections compact and efficient, but attention remains a mismatch: its K/V cache is created online and is typically processed by a separate higher-precision engine. Compressing this cache alone does not resolve the mismatch. To execute attention with the same lookup-table machinery as ternary projections, values accumulated in one reduction must retain a compatible representation and scale. This requirement also differs for keys and values during causal decoding, because newly generated values may belong to an unfinished cache block. This work develops a unified lookup-table inference approach for ternary LLMs. It stores runtime K/V states as scaled multi-plane signed digits organized around the reduction structure of attention. The resulting digit planes are consumed directly by activation-derived tables, avoiding dense K/V materialization between cache storage and attention computation. The design combines online K/V formation, bounded handling of incomplete value blocks, and a shared multi-stream datapath for Linear projections and attention. A constraint-guided search selects the representation and execution policy for a target quality--efficiency trade-off. Experiments on native and post-training ternary models validate the approach across cache capacity, model quality, and hardware efficiency.

↑