发表机构
PayPal Inc.(贝宝公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型API缓存成本,发现现有方法不足。提出缓存感知提示压缩(CAPC),结合查询无关压缩与缓存控制。经实验验证,CAPC在多种配置下成本最低,在生产工作负载中表现出色,验证了交叉模型负投资回报率预测。
AI 中文摘要
生产中的大语言模型部署结合了两种降低成本的方法:提示缓存(对重复使用的令牌前缀收取折扣率)和提示压缩(发送的令牌更少)。压缩文献采用查询感知方法,每个查询生成不同的压缩前缀,每次调用都会机械地使严格前缀缓存无效。我们在Anthropic的Sonnet 4.6 API上对这种成本进行了实证研究,发现缓存远非文献中假设的rho=1.0理想情况:Sonnet的缓存具有双层架构,在3500个令牌附近有一个明显的阈值,低于该阈值,在30次调用会话中命中率稳定在rho~0.83。我们的成本模型预测并经实验证实,在实际的rho下,查询感知压缩在高压缩率(r>=6)时优于朴素缓存。我们提出了缓存感知提示压缩(CAPC),将与查询无关的压缩与显式缓存控制以及防止过度压缩将缓存前缀推入热层的层保留比率界限相结合。在LongBench-v2的16/16配置中,CAPC是最便宜的策略,与仅使用缓存相比平均节省49%,与查询感知压缩相比节省64%,与普通方法相比节省90%,质量在未压缩基线的0.05以内。我们在三个生产工作负载上验证了CAPC:一个使用94k令牌模式前缀的企业工具助手(在r=3时成本降低51.7%);一个跨两个代码库的图形化知识图谱RAG管道(在FastAPI上是缓存所有内容的9.3倍,在httpx上是2.4倍);以及公共tau-bench零售基准(50个任务),其中CAPC是四种策略中最便宜的,奖励与普通方法完全相同(均为36/50,p=1.00),而查询感知压缩比普通方法贵40.1%——这是交叉模型在公共基准上的负投资回报率预测的首次生产验证。
英文摘要
Production LLM deployments combine two cost-reduction primitives: prompt caching (a discounted rate for re-used token prefixes) and prompt compression (fewer tokens sent). The compression literature has standardized on query-aware methods that produce a different compressed prefix per query, mechanically invalidating the prefix-strict cache on every call. We characterize this cost empirically on Anthropic's Sonnet 4.6 API and find caching is far from the rho=1.0 ideal the literature assumes: Sonnet's cache has a two-tier architecture with a sharp threshold near 3,500 tokens, below which the hit rate plateaus at rho~0.83 across 30-call sessions. Our cost model predicts, and experiments confirm, that under realistic rho, query-aware compression beats naive caching at high compression ratios (r>=6). We propose Cache-Aware Prompt Compression (CAPC), pairing query-agnostic compression with explicit cache_control plus a tier-preserving ratio bound that prevents over-compression from pushing the cached prefix into the hot tier. CAPC is the cheapest strategy in 16/16 configurations on LongBench-v2, with mean savings of 49% over cache-only, 64% over query-aware compression, and 90% over vanilla, at quality within 0.05 of the uncompressed baseline. We validate CAPC on three production workloads: an enterprise tool-using assistant with a 94k-token schema prefix (51.7% cost reduction at r=3); a graphify knowledge-graph RAG pipeline across two codebases (9.3x vs cache-all on FastAPI, 2.4x on httpx); and the public tau-bench retail benchmark (50 tasks), where CAPC is the cheapest of four strategies with reward exactly equal to vanilla (both 36/50, p=1.00) while query-aware compression is the most expensive at +40.1% over vanilla -- the first production confirmation of the crossover model's negative-ROI prediction on a public benchmark.
Comments28