发表机构
Independent Researcher, Milpitas, CA 95035 USA
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究发现大语言模型服务中,前缀缓存会引发智能体工具使用轨迹差异,且量化程度越高差异越显著,缓存状态缺失导致实际不可复现,同时发布了相关测试资源。
AI 中文摘要
前缀缓存是指服务引擎在不同请求间复用共享提示前缀的键值张量,已在主流开源栈中默认启用,被视为透明优化。本文研究其在可复现性上的代价,发现该代价随权重量化程度急剧上升。在固定模型、解码参数、随机种子、请求顺序,且以批量大小1串行发出每个请求的前提下,我们在两个引擎、四种权重格式下,分别运行了启用和禁用缓存的80集多轮智能体工具使用工作负载。启用缓存后,智能体轨迹在16位精度下有36.2%的集数发生变化,在4位精度下则有75.0%的集数发生变化,该梯度在受控缓存配置下可重复测量。禁用缓存时,所有配置下重复执行均为比特级一致,800集数中无差异,这将其他非确定性来源的比例限制在0.5%。启用缓存的重复运行确实会产生差异,三项实验定位了原因:单一服务器级提示缓存设置使运行间差异扩大了37.5个百分点,执行顺序仅在该设置激活时起作用,恢复缓存状态后,缓存路径和重新计算路径在40项中各有40项可复现,但彼此间仍有14项不同。给定缓存状态,缓存服务是确定性的,但实际上不可复现,因为该状态未包含在请求中且默认不重置。单轮桥接实验显示,差异会影响任务结果,但不会改变总体准确率。我们发布了测试框架、日志和分析流程。
英文摘要
Prefix caching, in which a serving engine reuses the key and value tensors of a shared prompt prefix across requests, is enabled by default in the major open-source stacks and treated as a transparent optimization. We measure what it costs in reproducibility, and find that the cost rises sharply with weight quantization. Holding the model, decoding parameters, seed, and request order fixed, and issuing every request serially at batch size one, we ran an eighty-episode multi-turn agentic tool-use workload with caching enabled and disabled across two engines and four weight formats. Enabling the cache changed the agent's trajectory on 36.2 percent of episodes at 16-bit precision and on 75.0 percent at four-bit, a gradient that survives re-measurement under a controlled cache configuration. With caching disabled, repeated execution was bit-identical in every configuration, 0 of 800 episodes, which bounds other sources of nondeterminism at 0.5 percent. Repeated cache-enabled runs did diverge, and three experiments locate the cause: a single server-level prompt-cache setting moves run-to-run divergence by 37.5 percentage points, execution order acts only while that setting is active, and restoring cache state makes the cached and recompute paths each reproduce on 40 of 40 items while still differing from each other on 14. Cached serving is deterministic given cache state, and irreproducible in practice because that state is absent from the request and never reset by default. A single-turn bridge shows the divergence reaching task outcomes without shifting aggregate accuracy. We release the harness, logs, and analysis pipeline.
CommentsSubmitted to IEEE Access