超越固定故障模型:比较OpenStack中基于LLM与基于规则的故障注入
Beyond Fixed Fault Models: Comparing LLM-Based and Rule-Based Fault Injection in OpenStack
浏览论文内容
中文总结 AI 辅助
本研究比较LLM与规则故障注入在OpenStack中的效果,发现LLM扩展了故障行为覆盖,但需受控生成和验证才能实际应用。
中文摘要 AI 辅助
软件故障注入(SFI)通过引入软件缺陷并观察其表现来支持云系统的测试。基于规则的注入器(如ProFIPy)提供受控且可复现的源代码级变异,但需要手动编码故障模式。大型语言模型(LLMs)通过生成上下文相关的软件故障提供了一种数据驱动的替代方案。我们比较了两种代码LLM(Qwen2.5-Coder和DeepSeek-Coder)与ProFIPy在OpenStack的Nova和Cinder服务中的表现。在共享注入目标上,激活率和可观察故障率相当,但操作特征不同:LLM生成的故障在Nova上产生更多灾难性结果,而ProFIPy产生更多静默和多组件效应。采样的LLM输出在故障表现方式上也有所不同,但在传播范围上表现出更大的一致性。这些发现表明,基于LLM的故障注入扩展了固定故障模型的行为覆盖范围,但并未确立普遍优越性,实际采用仍需受控生成、运行时验证、系统级预言机以及可复现的实验来源。
英文摘要
Software Fault Injection (SFI) supports testing of cloud systems by introducing software defects and observing their manifestation. Rule-based injectors such as ProFIPy provide controlled and reproducible source-level mutations but require fault patterns to be encoded manually. Large Language Models (LLMs) offer a data-driven alternative by generating context-dependent software faults. We compare two code LLMs, Qwen2.5-Coder and DeepSeek-Coder, with ProFIPy in OpenStack's Nova and Cinder services. On shared injection targets, activation and observable-failure rates are comparable, but operational profiles differ: LLM-generated faults produce more Catastrophic outcomes on Nova, whereas ProFIPy produces more Silent and Multi-component effects. The sampled LLM outputs also differ in how they manifest failure, while showing greater agreement in their propagation scope. These findings show that LLM-based fault injection extends the behavioral coverage of fixed fault models without establishing general superiority, and that practical adoption still requires controlled generation, runtime validation, system-level oracles, and reproducible experimental provenance.
发表机构
- University of Naples Federico II(那不勒斯费德里科二世大学)
- IMT School for Advanced Studies Lucca(卢卡高等研究学院)
- University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)
机构由 AI 辅助整理,请以论文原文为准。