arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

冷却地雷:LLM 网关上的跨租户干扰攻击

Cooldown Landmines: Cross-Tenant Interference Attacks on LLM Gateways

Yudong Gao, Linghan Chen, Wenhan Wu, Quan Shi, Xutao Mao, Mia Zhou, Junjian Li, Xiaolong Liu, Jiyao Wang, Mingyu Guo, Honglong Chen

arXiv 2610.05089首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; University of Adelaide; Wuhan University; National University of Singapore; City University of Hong Kong; University of North Carolina at Chapel Hill; Geely; Hunan University; ETH Zurich; China University of Petroleum(香港科技大学; 阿德莱德大学; 武汉大学; 新加坡国立大学; 香港城市大学; 北卡罗来纳大学教堂山分校; 吉利; 湖南大学; 苏黎世联邦理工学院; 中国石油大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示LLM网关因共享冷却记录导致的跨租户干扰攻击,提出来源检查与租户范围冷却,实验证明可显著降低受害者回退。

AI 中文摘要

LLM 网关在共享模型部署和冷却记录的同时强制执行独立的租户配额,这些冷却记录会临时排除故障的后端。然而,一个租户的请求失败可能会更新这些共享记录,并限制其他租户访问可服务的部署。我们识别了两种利用 LiteLLM 中这一漏洞的攻击。第一种利用在密钥的每分钟请求数(RPM)限制下被拒绝的请求:调用者提供的已知注册部署标识符会到达故障处理,使得两个被拒绝的请求能够将另一个租户重定向到回退,而无需这些请求产生任何上游调用。第二种利用已接纳的流量创建在后台容量恢复后仍持续存在的冷却记录。为了解决这些故障,我们设计了一种来源检查,阻止来自密钥 RPM 拒绝的部署更新,以及一种租户范围的冷却,保留触发租户的回退,同时保留后端故障的共享记录。使用认证代理和自托管 vLLM 进行的实验表明,这些攻击可以在部署保持可服务的同时强制回退或拒绝服务。在五对四工作进程的运行中,租户范围化将恢复后的受害者回退从 54/60 降至 0/60,同时增加了对已耗尽共享配额的尝试。这些发现表明,租户隔离必须涵盖请求接纳和管理共享部署可用性的故障处理。

英文摘要

LLM gateways enforce separate tenant quotas while sharing model deployments and cooldown records that temporarily exclude failing backends. However, a tenant's request failure can update these shared records and restrict other tenants' access to serviceable deployments. We identify two attacks that exploit this gap in LiteLLM. The first uses requests rejected at the key's requests-per-minute (RPM) limit: caller-supplied identifiers for known registered deployments reach failure handling, allowing two rejected requests to redirect another tenant to fallback with zero upstream calls from those requests. The second uses admitted traffic to create cooldown records that persist after backend capacity recovers. To address these failures, we design an origin check that blocks deployment updates from key RPM rejections and tenant-scoped cooldown that preserves the triggering tenant's back-off while retaining shared records for backend faults. Experiments with authenticated proxies and self-hosted vLLM demonstrate that the attacks can force fallback or denial while deployments remain serviceable. Across five paired four-worker runs, tenant scoping reduces victim fallback after recovery from 54/60 to 0/60, while increasing attempts against exhausted shared quotas. These findings show that tenant isolation must cover both request admission and the failure handling that governs shared deployment availability.

Comments24 pages, 7 figures. Main text in ICLR format; includes appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑