arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ContinuityBench:多提供商大语言模型路由中有状态故障转移的基准测试与系统研究

ContinuityBench: A Benchmark and Systems Study of Stateful Failover in Multi-Provider LLM Routing

Vishal Pandey, Gopal Singh

arXiv 2607.15899首次发表:更新:

发表机构

Metriqual(梅特里夸尔)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多提供商LLM路由中状态故障转移问题,提出有状态代理架构及历史转发策略,引入新指标,通过continuity - bench评估,实证表明该架构CPR达99.20%,刻画延迟分布,为构建多模型推理系统提供基础。

AI 中文摘要

在生产环境中的大语言模型(LLM)部署中,高API可用性并不等同于对话连续性。当主提供商出现故障或严格的速率限制时,简单的无状态故障转移机制虽能维持正常运行时间,但会悄无声息地丢弃对话历史,严重影响用户体验。为严格量化和解决此故障模式,我们引入了两个新指标:连续性保留率(CPR)和连续性延迟开销(CLO)。我们提出了一种有状态的多提供商代理架构,利用历史转发策略在故障转移事件中跨异构LLM端点无缝重建对话状态。此外,我们发布了continuity - bench(此https URL),一个开放评估工具,用于在高并发提供商故障条件下压力测试上下文保留。我们的实证评估(750次故障转移事件)表明,与标准无状态架构近乎0%的保留率相比,我们的有状态代理实现了99.20%的CPR[95%置信区间:98.27%,99.63%],能将深入的对话上下文干净地转移到备用提供商。最后,我们刻画了故障转移延迟分布,确定了带抖动的异步指数退避对于防止针对严格限制的备用API的级联重试风暴的关键必要性。我们的结果为构建强大的、保留状态的多模型推理系统提供了原则性基础。

英文摘要

In production large language model (LLM) deployments, high API availability guarantees do not equate to conversational continuity. When a primary provider experiences an outage or strict rate-limiting, naive stateless failover mechanisms successfully maintain uptime but silently discard conversation history, severely disrupting the user experience. To rigorously quantify and resolve this failure mode, we introduce two novel metrics: Continuity Preservation Rate (CPR) and Continuity Latency Overhead (CLO). We propose a stateful, multi-provider proxy architecture utilizing a History-Forwarding strategy to seamlessly reconstruct conversational state across heterogeneous LLM endpoints during failover events. Furthermore, we release continuity-bench, https://github.com/Vishal-sys-code/continuity-bench, an open evaluation harness designed to stress-test context preservation under high-concurrency provider failure conditions. Our empirical evaluation ($N=750$ failover events) demonstrates that our stateful proxy achieves a 99.20\% CPR [95\% CI: 98.27\%, 99.63\%], cleanly transferring deep conversational context to fallback providers, compared to a near-0\% preservation rate for standard stateless architectures. Finally, we characterize failover latency distributions, identifying the critical necessity of asynchronous exponential backoff with jitter to prevent cascading retry storms against strict-limit fallback APIs. Our results provide a principled foundation for building robust, state-preserving multi-model inference systems.

Comments16 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑