AI系统的可靠性工程:挑战、方法与方向
Reliability Engineering for AI Systems: Challenges, Methods, and Directions
浏览论文内容
中文总结 AI 辅助
本文借鉴传统可靠性工程方法,提出四级故障诊断框架及TEVV、FRACAS等工具,用于系统化评估和提升AI系统可靠性,并通过三个案例展示其应用。
中文摘要 AI 辅助
AI可靠性关注的是AI系统在规定时段内、在规定运行条件下,凭借既定证据,能否可靠地执行其预期功能。随着这些系统变得更加自主,该功能已不仅限于产生正确输出。检索、记忆、工具使用、权限、人工监督以及系统间的交互必须一致且安全地运行,对于生成式系统,产生输出的推理过程也必须如此。平均基准准确率衡量的是能力;它并未量化这一更广泛的可靠性主张。本文将对已建立的可靠性工程方法进行改编,从故障定义和运行包络,到FMEA、加速测试、现场监测和可靠性增长,以适用于AI系统。一个四级诊断框架将故障分类为组件故障、运行回路故障、智能体行为故障或网络与治理故障。测试、评估、验证和确认(TEVV)、序贯监测和FRACAS创建并刷新证据。SMART为测量、分析、评估和测试规划提供统计指导;NIST AI风险管理框架为治理、评估、监测和缓解提供组织指导。三个案例说明了该程序:对卷积神经网络的对抗性测试、感知误差传播以及自动驾驶车辆脱离。已建立的可靠性工程提供了一个可用的基础;随着这些系统自我演进,仍需要新的测量和安全护栏。
英文摘要
AI reliability concerns whether an AI system performs its intended function dependably over a stated period and under stated operating conditions, with stated evidence. As these systems become more autonomous, that function includes more than a correct output. Retrieval, memory, tool use, permissions, human oversight, and interactions among systems must operate consistently and safely, and, for generative systems, so must the reasoning process that produces the output. Average benchmark accuracy measures capability; it does not quantify this broader reliability claim. This paper adapts established reliability engineering methods, from failure definitions and operational envelopes to FMEA, accelerated testing, field monitoring, and reliability growth, to AI systems. A four-level diagnostic framework classifies failures as component, operational-loop, agentic-conduct, or network and governance failures. Test, evaluation, verification, and validation (TEVV), sequential monitoring, and FRACAS create and refresh evidence. SMART provides statistical guidance for measurement, analysis, assessment, and test planning; the NIST AI Risk Management Framework provides organizational guidance for governance, evaluation, monitoring, and mitigation. Three cases illustrate the program: adversarial testing of a convolutional neural network, perception-error propagation, and autonomous-vehicle disengagements. Established reliability engineering provides a usable foundation; new measurements and safety guardrails are still needed as these systems are self-evolving.
发表机构
- Arizona State University(亚利桑那州立大学)
- Virginia Tech(弗吉尼亚理工大学)
- City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。