arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向模拟存算一体系统上抗噪声大语言模型推理的选择性键值缓存保护

Selective KV Cache Protection for Noise-Resilient LLM Inference on Analog Compute-In-Memory Systems

Yuannuo Feng, Wenyong Zhou, Yuang Ma, Yizhe Chen, Wenshuai Yao, Yuxin Xie, Ngai Wong, Wang Kang

arXiv 2607.29076首次发表:更新:

AI 中文总结

针对模拟存算一体系统上LLM推理的KV缓存易受硬件噪声影响的问题,本文首次开展相关系统研究,提出分层token保护策略,大幅降低噪声下的困惑度并提升编程行利用率。

AI 中文摘要

模拟存算一体(CIM)阵列已成为大语言模型(LLM)高效推理的有潜力载体,尤其适用于线性层的权重驻留计算。然而,将模拟CIM扩展到注意力机制时会引入一个基础挑战:键值(KV)缓存操作需要重复的原位权重更新,这与权重驻留范式不匹配,使动态计算面临严重的硬件噪声问题,该关键问题在很大程度上仍未被探索。本文首次对模拟CIM阵列上的动态注意力计算进行系统研究,发现初始token和近期token对硬件噪声表现出不成比例的脆弱性。受此token级见解启发,我们提出分层token保护策略:将初始token(sink tokens)和滑动近期token窗口保留在更高精度的数字路径上,同时在模拟CIM上处理大部分KV缓存。协同设计的调度器结合模拟编程、所有权转换和批量矩阵-矩阵乘法(MVM)块形成,以限制数字开销。对9种LLM的评估显示,我们的方法将模拟噪声下的平均困惑度从33.91降至11.95,接近干净基线的11.06,同时将动态KV编程行利用率从23.1%提高到91.2%。

英文摘要

Analog compute-in-memory (CIM) arrays have emerged as a promising substrate for energy-efficient LLM inference, particularly for weight-stationary computations in linear layers. However, extending analog CIM to attention mechanisms introduces a fundamental challenge: KV cache operations demand repeated in-situ weight updates, and the resulting mismatch with the weight-stationary paradigm exposes dynamic computations to significant hardware noise, a critical problem that remains largely unexplored. In this paper, we present the first systematic study of dynamic attention computation on analog CIM arrays, revealing that initial and recent tokens exhibit disproportionate vulnerability to hardware noise. Motivated by this token-level insight, we propose a hierarchical token protection strategy that keeps sink tokens and a sliding recent-token window on a higher-precision digital path while processing the bulk KV cache on analog CIM. A co-designed scheduler combines analog programming, ownership transition, and bulk-MVM tile formation to bound digital overhead. Evaluations on nine LLMs show that our approach lowers average perplexity under analog noise from 33.91 to 11.95, close to the clean baseline of 11.06, while improving dynamic-KV programming-row utilization from 23.1\% to 91.2\%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑