AI 中文总结
该研究提出KeyPooling测量方法,发现LLM API中继服务默认未将客户绑定到上游凭证,存在跨客户缓存读取漏洞,并提出防御契约及优化方案。
AI 中文摘要
大型语言模型(LLM)API中继服务会分别对客户进行身份验证,但通常会通过共享的提供商凭证转发请求。提供商将提示缓存的作用范围限定为上游主体和命名空间,因此映射到同一缓存身份的中继客户可以观察彼此的缓存状态。先前的研究显示部分端点存在缓存共享现象,但未明确是哪个凭证、池、适配器或嵌套跳点控制了最终身份。我们提出KeyPooling这一测量方法,该方法可追踪客户在缓存查找和写入过程中的身份,验证运行时转换,并一次测试一个预测的身份组件。在连接OpenAI和Anthropic的5个开源网关中,默认情况下没有一个网关将客户绑定到上游凭证;在共享凭证下,5个网关均暴露了针对这两家提供商的跨客户缓存读取。主体与命名空间拆分、池关联以及适配器和嵌套中继的对比,明确了控制转换的位置。在一个与结果无关的每周OpenRouter框架中,测试覆盖了80.5%的合格令牌量,发现28个标签中有12个存在跨账户读取,占33.7%的量。在一条生产路由上,一项受控程序在未访问目标的情况下恢复了8个连续的目标位置。更广泛的测试确定了缓存粒度、路由、速率限制、归因和预算是逐令牌恢复的条件,而非安全控制。我们提出一项防御契约:每个客户必须进入提供商强制的域,或源自已认证身份的命名空间必须在每次最终缓存查找和写入中保留。将此拆分置于可重用公共前缀之后,在1.7-2.5%的成本增加下保留了大部分建模的重用性。
英文摘要
Large language model (LLM) API relays authenticate customers separately but often forward requests through shared provider credentials. Providers scope prompt caches to upstream principals and namespaces, so relay customers mapped to one cache identity can observe each other's cache state. Prior work showed cache sharing at selected endpoints but did not identify which credential, pool, adapter, or nested hop controls the finalidentity. We present KeyPooling, a measurement method that traces customer identity through cache lookup and write, verifies runtime transformations, and tests one predicted identity component at a time. Across five open-source gateways connected to OpenAI and Anthropic, none bound customers to upstream credentials by default; under a shared credential, all five exposed cross-customer cache reads for both providers. Principal and namespace splits, pool associations, and adapter and nested-relay contrasts localized the controlling transformations. In an outcome-independent weekly OpenRouter frame, tests covered 80.5% of eligible token volume and found cross-account reads for 12 of 28 labels carrying 33.7% of volume. On one production route, a controlled procedure recovered eight consecutive target positions without target access. Broader tests identify cache granularity, routing, rate limits, attribution, and budget as conditions for token-by-token recovery, not security controls. We derive a defense contract: every customer must enter a provider-enforced domain, or a namespace derived from authenticated identity must survive every final cache lookup and write. Placing this split after reusable public prefixes preserved most modeled reuse at a 1.7-2.5% cost increase.