AI 中文总结
提出AgentChaos框架,在不修改源代码的情况下于LLM API共享层注入故障,评估显示智能体系统鲁棒性取决于自身实现,现有故障诊断方法仍有提升空间。
AI 中文摘要
智能体系统的每一次响应都依赖大语言模型(LLM)API,但这些API可能返回服务器错误、截断的响应或损坏的内容,这些问题会在下游智能体中传播并导致任务失败。评估这些故障下的鲁棒性对可靠部署至关重要。现有的故障注入方法多为离线方式,需要修改源代码,且无法修改特定响应字段;全面评估还需要系统的故障分类,因为不同故障类型对下游智能体的影响存在差异。我们提出AgentChaos,这是一个用于受控、运行时、非侵入式LLM API故障注入的混沌工程框架。由于所有智能体系统都通过同一HTTP接口访问LLM,我们在该共享层注入故障,无需修改源代码。我们定义了内容和工具调用字段上的崩溃故障、遗漏故障和值故障,在运行时拦截并修改LLM API响应,同时验证每个故障是否被触发,以过滤未触发的任务,避免低估故障影响。在65种故障配置下,对多个智能体系统、基准测试和骨干LLM的评估显示,所有系统在故障注入下均出现性能下降,pass@1下降幅度最高达50个百分点;不同模型间的性能排名一致,表明鲁棒性取决于系统实现而非模型能力。现有故障诊断方法在故障类型识别上准确率低于53%,在故障步骤识别上低于56%,仍有改进空间。我们还为智能体系统开发者揭示了实用发现。
英文摘要
Agent systems rely on LLM APIs for every response, but these APIs can return server errors, truncated responses, or corrupted content that propagates through downstream agents and causes task failure. Evaluating robustness under these faults is crucial for reliable deployment. Existing fault injection methods are offline, require source code modification, or cannot modify specific response fields. A comprehensive evaluation also requires a systematic fault taxonomy because different fault types affect downstream agents differently. We propose AgentChaos, a chaos engineering framework for controlled, runtime, non-intrusive LLM API fault injection. Since all agent systems access LLMs through the same HTTP interface, we inject faults at this shared layer without modifying source code. We define crash, omission, and value faults on content and tool call fields, intercept and modify LLM API responses at runtime, and verify whether each fault is triggered to filter untriggered tasks and avoid underestimating fault impact. Evaluations across agent systems, benchmarks, and backbone LLMs under 65 fault configurations show that all systems degrade under fault injection, with pass@1 dropping by up to 50 percentage points. The ranking is consistent across models, suggesting that robustness depends on system implementation rather than model capability. Existing fault diagnosis methods achieve below 53% accuracy on fault type and below 56% on fault step, leaving room for improvement. We further reveal practical findings for agent system developers.
CommentsAccepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026)