发表机构
Jeju National University; Soongsil University(济州国立大学; 崇实大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示HQQ缓存量化中共享停止决策导致批量无关问题改变答案,提出固定迭代预算或请求局部停止以消除依赖,并强调审计需覆盖停止决策。
AI 中文摘要
语言模型系统将问题批量处理以提高吞吐量,但无关的问题不应改变目标问题的答案,因为其输入和数值执行是固定的。我们研究了键值缓存的压缩,该缓存存储生成过程中复用的注意力表示。在请求局部组中,Transformer的半二次量化(HQQ)后端分别更新压缩参数,但使用共享的平均误差来决定所有更新何时停止。仅替换与目标问题批量处理的问题,在两个模型的384次测试比较中,有170次改变了四比特HQQ的答案。重放另一执行过程的更新次数,在每次改变的对中(双向)都能重现其完整答案和缓存指纹。在FP32中计算停止平均值减少了缓存差异,但保留了答案变化。原生HQQ在八个算术对中也改变了已确认的数值正确性。固定迭代次数和请求局部停止在匹配控制下消除了观察到的伴随依赖。请求局部停止在张量级别上仍对合成填充变化敏感。修复原始迭代预算无需调整即可消除此决策路径。两种修复均无既定质量优势,且自然重新批处理仍会改变答案。请求独立性审计必须涵盖停止决策以及量化组。
英文摘要
Language-model systems batch questions for throughput, but unrelated questions should not change a target's answer when its input and numerical execution are fixed. We study compression of the key and value cache, which stores attention representations reused during generation. With request-local groups, Transformers' Half-Quadratic Quantization (HQQ) backend updates compression parameters separately but uses a shared average error to decide when all updates stop. Replacing only the question batched with the target changes four-bit HQQ answers in 170/384 test comparisons across two models. Replaying the other execution's update counts reproduces its complete answer and cache fingerprints in every changed pair, in both directions. Computing the stopping mean in FP32 reduces cache differences but leaves answer changes. Native HQQ also changes confirmed numerical correctness in eight arithmetic pairs. Fixed iterations and request-local stopping remove observed companion dependence under matched controls. Request-local stopping remains sensitive to synthetic padding changes at the tensor level. Fixing the original iteration budget removes this decision path without tuning. Neither repair has an established quality advantage, and natural rebatching still changes answers. Request-independence audits must cover stopping decisions as well as quantization groups.