arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

保持缓存热度是值得的:智能工作负载的保活经济学

Keeping the Cache Warm Pays: Keepalive Economics for Agentic Workloads

Maxim Khailo

arXiv 2607.19214首次发表:更新:

AI 中文总结

研究智能工作负载中缓存前缀易被逐出致成本增加的问题,提出客户端保活方法,通过定时器重放前缀保持缓存热度,降低请求成本,还探讨了ping频率及保活策略的经济性,推导了运营商策略。

AI 中文摘要

前沿大语言模型(LLM)提供商缓存提示的已处理前缀,后续共享该前缀的请求只需支付约10%的输入价格,并跳过大部分预填充延迟。但智能工作负载会破坏这一好处:智能体发送请求、运行工具或等待批准数分钟,后续请求发送时缓存前缀已被逐出,需再次支付全部预填充费用。客户端保活可防止这种情况,通过定时器在暂停期间重放前缀,能在空闲基线被逐出的间隙保持前缀热度,将暂停后请求成本降低多达12.5倍。战略问题是ping频率,经济选择是在提供商TTL安全范围内的最大间隔,如Anthropic的5分钟TTL下约4分钟,而非30秒惯例,该策略在空闲时与重新预填充达到收支平衡。由于好处真实且仅受用户账单限制,合理采用即普遍采用;且缓存驻留按读取计费,保活饱和层使LRU逐出无物可排。我们认为这种外部性将促使提供商直接计量缓存驻留,已有一家这样做了。我们推导了运营商在此之前的策略。

英文摘要

Frontier LLM providers cache a prompt's processed prefix so that a follow-up request sharing it pays ~10% of the input price and skips most of the prefill latency. Agentic workloads systematically destroy this benefit: the agent sends a request, runs a tool or waits for approval for minutes, and by the time the follow-up is sent the cached prefix has been evicted, so the agent pays the full prefill again. A client-side keepalive, replaying the prefix on a timer during the pause, prevents this, and it is individually rational: across Anthropic, OpenAI, Google, and DeepSeek we show that a keepalive holds the prefix warm through gaps where idle baselines are evicted, cutting the post-pause request cost by up to 12.5x. The strategic question is the ping frequency, and it has a clean answer: keepalive cost falls monotonically in the interval, so the economical choice is the largest interval safely under the provider's TTL, about 4 minutes at Anthropic's 5-minute TTL rather than the 30-second convention, and the strategy breaks even against a re-prefill at idle ~tau(w/r - 1) (~46 min for Anthropic, ~36 min for OpenAI and DeepSeek). Because the benefit is real and bounded only by each user's own bill, rational adoption is universal adoption; and since cache residency is priced per read rather than per token-hour, a keepalive-saturated tier gives LRU eviction nothing to rank. We argue this externality will push providers to meter cache residency directly, and one already does. We derive the operator's policy until then.

Comments5 pages, 1 figure, 3 tables. Measurement data and harness: https://github.com/mempko/pi (scripts/cache-research)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑