AI 中文总结
本文提出使用确定性黄金比例Weyl抖动替代随机舍入来量化Mamba风格模型的循环状态缓存,在多种模型和长序列生成中始终更接近全精度结果,且无需额外成本。
AI 中文摘要
Mamba风格和混合语言模型将其过去的信息压缩为固定大小的循环状态,并在每个生成的token处重写该状态。以低精度存储此状态可节省内存带宽,但每次舍入误差都会反馈到下一次更新中,并可能在长序列生成过程中累积。生产系统会随机舍入状态;我们研究此类缓存应采用哪种舍入规则。我们发现,一种确定性的黄金比例Weyl抖动(无需随机数)在纯模型和混合模型、不同存储格式以及长解码视野下,始终比随机舍入更接近全精度模型的量化结果,且无需额外成本。舍入到最近邻的行为则不同:由于它会丢弃小的更新,其误差持续增长,因此在短评估中可能表现最佳,但在长序列生成中却远远落后。差异分析解释了这一排序,我们还记录了会悄然消除该优势的实现陷阱。
英文摘要
Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this state in low precision saves memory bandwidth, but every rounding error is fed back into the next update and can accumulate over long generations. Production systems round the state stochastically; we ask which rounding rule such caches should use. We find that a deterministic golden-ratio Weyl dither, which needs no random numbers, consistently brings the quantized model closer to the full-precision one than stochastic rounding, across pure and hybrid models, storage formats, and long decoding horizons, at no extra cost. Round-to-nearest behaves differently: because it discards small updates, its error keeps growing, so it can look best in short evaluations yet falls far behind over long generations. A discrepancy analysis explains this ordering, and we document implementation pitfalls that silently remove the benefit.