arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对HAProxy中后端故障隔离的LLM引导控制平面策略进行基准测试

Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy

Aman Chauhan, Vishnu Pendyala

arXiv 2608.10532首次发表:更新:

发表机构

San José State University(圣何塞州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文测试了不同规模和架构的LLM在HAProxy后端故障隔离中的策略效果,发现约3B有效参数的能力阈值,达标后可显著降低5xx错误,但会增加延迟与令牌开销,高效方案为阈值以上模型的非推理模式。

AI 中文摘要

静态负载均衡器无法缓解处于降级状态而非完全下线的后端:轮询和最小连接数算法会持续将流量路由到返回HTTP 500错误的服务器,直到运维人员介入。本文研究大型语言模型(LLM)是否能替代静态路由策略,每10秒读取HAProxy和Prometheus遥测数据,并通过受管控的HAProxy数据平面API调用隔离故障服务器。在一个可复现的基准测试中,异构集群里约三分之一的服务器内置了持续性结构故障,本文对5个系列的15个开源权重模型(总参数规模0.35B至35B;包含密集型、混合专家型及高效稀疏型架构)、推理模式、3至9个后端的集群规模以及2种路由算法进行了遍历测试,总计240次运行。研究发现,有效参数规模存在约3B的能力阈值:低于该阈值时,LLM策略通常不可靠,有时甚至比无策略更差;高于该阈值时,无论架构如何,所有模型均能使客户端感知到的5xx错误较静态基准降低近88%。该阈值是近似的:Gemma 4 E2B以2B有效参数即可达标,而密集型的3B Granite 4.0 Micro则无法达标。可用性提升存在代价:引流会将负载集中到幸存服务器,使尾部延迟膨胀2.6至2.8倍;启用推理会使令牌开销增加约10倍,超出控制间隔并降低有效性。高效运行点为阈值以上模型的最便宜非推理模式,且被确定性管控措施包裹。

英文摘要

Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server returning HTTP 500s until an operator intervenes. We ask whether a Large Language Model can replace the static routing policy itself, reading HAProxy and Prometheus telemetry every 10 seconds and isolating faulty servers through guardrailed calls to the HAProxy Data Plane API. On a reproducible benchmark with a persistent structural fault built into roughly one-third of a heterogeneous fleet, we sweep 15 open-weight models across five families (0.35B to 35B total parameters; dense, mixture-of-experts, and efficient-sparse architectures), reasoning modes, fleet scales of 3 to 9 backends, and two routing algorithms, totaling 240 runs. We find a capability threshold near 3B active parameters. Below it, LLM policies are typically unreliable and sometimes worse than no policy; above it, every model, regardless of architecture, saturates near an 88% reduction in client-perceived 5xx errors over the static baseline. The threshold is approximate: Gemma 4 E2B clears it with 2B active parameters, while the dense 3B Granite 4.0 Micro does not. The availability gain has costs. Draining concentrates load onto surviving servers, inflating tail latency 2.6 to 2.8 times, and enabling reasoning multiplies token spend roughly tenfold, overrunning the control interval and degrading effectiveness. The efficient operating point is a supra-threshold model in its cheapest non-reasoning mode, wrapped inside deterministic guardrails.

Comments43 pages, 6 figures, 15 tables. Submitted to Journal of Network and Computer Applications (Elsevier)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑