发表机构
Metriqual(梅特里夸尔)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多供应商大语言模型服务基础设施故障模式,提出FailureAtlas双轴分类法,依起源层和可检测性分类,通过公共报告和测试填充,发现最严重故障是隐蔽的,如并发竞争致历史丢失和流索引冲突,还给出重现脚本。
AI 中文摘要
多供应商大语言模型网关反向代理在基础模型API之间路由、负载均衡和限速请求,已成为关键的生产基础设施。然而,这一架构层特有的故障模式仍未记录,分散在问题跟踪器和事后分析中,没有统一框架。我们引入FailureAtlas,这是一种双轴分类法,按故障的“起源层”(网络/传输、流/协议、状态/会话、模型行为、治理/成本)和“可检测性”(明显与隐蔽)进行分类。我们用五个来自公共错误报告和第一手压力测试的经过验证的目录条目填充此分类法,每个条目都伴有机械性根本原因分析。三个条目包括独立的重现脚本。我们的主要发现是,最具操作严重性的故障是“隐蔽的”:它们返回HTTP 200,通过每一项标准健康检查,并以需要语义级可观测性才能检测到的方式破坏应用程序状态。在CB评估活动中,首次直接发现了两个这样的隐蔽故障:一个导致历史记录丢失的并发竞争条件和一个破坏工具调用有效负载的流索引冲突。
英文摘要
Multi-provider LLM gateways reverse proxies that route, load-balance, and rate-limit requests across foundation-model APIs have become critical production infrastructure. Yet the failure modes specific to this architectural layer remain undocumented, scattered across issue trackers and post-mortems with no unifying framework. We introduce \fa{}, a two-axis taxonomy that classifies failures by their \emph{origin layer} (Network/Transport, Streaming/Protocol, State/Session, Model~Behavior, Governance/Cost) and their \emph{detectability} (Loud vs.\ Silent). We populate this taxonomy with five verified catalog entries sourced from public bug reports and first-hand stress testing, each accompanied by a mechanistic root-cause analysis. Three entries include standalone reproduction scripts. Our principal finding is that the most operationally severe failures are \emph{silent}: they return HTTP~200, pass every standard health check, and corrupt application state in ways that require semantic-level observability to detect. Two such silent failures a concurrency race condition causing history loss and a streaming index collision corrupting tool-call payloads were discovered first-hand during \cb{} evaluation campaigns.
CommentsSurvey Paper, 14 pages, 1 figure