发表机构
Boston University; University of Texas at Austin(波士顿大学; 德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究分析150起生产事故,提出含23种故障模式的复合AI系统分类法及韧性模式目录,实验显示断路器、质量门和隔离分别减少89%级联、捕获73%静默退化、减少64%爆炸半径,多模式采用使MTTR降低71%。
AI 中文摘要
可靠且安全地部署复合AI系统需要理解在组件边界处出现的故障模式,而非单个模型内部的故障。级联错误在组件边界间传播,静默质量退化规避了标准监控,协调失败导致由各自正确的部分产生错误的集体行为。我们分析了来自开源复合AI项目和匿名企业部署的150份生产事故报告,构建了一个包含23种故障模式的分类体系,这些模式分为五类:检索故障、生成故障、工具故障、编排故障和集成故障。针对每个类别,我们提出了通过受控故障注入实验测得的有效性的韧性模式。断路器将级联传播减少了89%,输出质量门在影响用户前捕获了73%的静默退化,组件隔离将爆炸半径减少了64%。采用我们目录中三个或更多韧性模式的系统,与无结构监控基线相比,平均恢复时间(MTTR)减少了71%。我们将事故分类和模式目录作为实践者资源发布。
英文摘要
Deploying compound AI systems reliably and safely requires understanding failure modes that emerge at component boundaries, not within individual models. Cascading errors propagate across component boundaries, silent quality degradation evades standard monitoring, and coordination failures yield incorrect collective behavior from individually correct parts. We analyze 150 production incident reports from open-source compound AI projects and anonymized enterprise deployments to construct a taxonomy of 23 failure modes organized into five categories: retrieval failures, generation failures, tool failures, orchestration failures, and integration failures. For each category, we propose resilience patterns with measured effectiveness from controlled fault injection experiments. Circuit breakers reduce cascade propagation by 89%, output quality gates catch 73% of silent degradation before user impact, and component isolation reduces blast radius by 64%. Systems implementing three or more resilience patterns from our catalog reduce mean-time-to-recovery (MTTR) by 71% compared to unstructured monitoring baselines. We release the incident taxonomy and pattern catalog as a practitioner resource.
CommentsAccepted at the AIWILD Workshop, ICML 2026. Camera-ready version