超越组件测试:验证智能体AI系统
Beyond Component Testing: Validating Agentic AI Systems
浏览论文内容
中文总结 AI 辅助
该综述综合257篇相关文献,构建五维分类法分析智能体AI系统验证问题,指出现有方法的缺口,提出面向生命周期的研究议程,主张需在上下文中验证智能体轨迹以实现可信部署。
中文摘要 AI 辅助
智能体AI系统通过结合规划、工具使用、记忆、交互与适应的多步骤轨迹执行动作,这种行为将验证实践延伸至组件测试及单次输入-输出评估之外,因为可接受的系统行为如今取决于决策随时间的展开方式以及在变化环境条件下的表现。本综述综合了257篇涵盖智能体评估、软件保障、网络物理系统、运行时监控及监管指南的论文,以明确智能体系统的验证问题。综述围绕涵盖行为、安全、时间、监管及多智能体问题的五维分类法展开,并用该分类法映射现有方法,揭示反复出现的覆盖缺口。分析显示,行为评估相对成熟,而时间有效性、运行时证据维护、监管可追溯性及开放式多智能体系统保障仍待发展。三项跨领域案例研究(医疗保健、工业运营、智能移动系统)基于综述文献记录的故障模式,说明五维分类法在安全关键场景中的应用。论文以面向生命周期的研究议程作结,该议程围绕受限自主权规范、对抗性轨迹生成、运行时监控及可审计证据结构展开,核心主张为:智能体AI的可信部署依赖于在上下文中验证轨迹,而非仅评估孤立组件。
英文摘要
Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior stretches validation practice beyond component testing and one-shot input-output evaluation, because acceptable system behavior now depends on how decisions unfold over time and under changing environmental conditions. This systematic mapping study synthesizes 262 papers spanning agent evaluation, software assurance, cyber-physical systems, runtime monitoring, and regulatory guidance in order to characterize the validation problem for agentic systems. The review is organized around a five-dimension taxonomy covering behavioral, safety, temporal, regulatory, and multi-agent concerns, and uses that taxonomy to map current approaches and identify dense and sparse pairings of approaches and dimensions. The resulting map shows that the literature concentrates on behavioral evaluation (105 of 262 papers) and is thinnest on temporal validity (14 papers); regulatory work rests mainly on assurance cases and regulatory analysis, largely from IEEE-indexed venues, and multi-agent work mainly on benchmarks. Three cross-domain case studies (medical care, industrial operations, smart-mobility systems) provide operational illustrations of how the five taxonomy dimensions recur in safety-critical settings, motivated by the failure patterns documented in the reviewed literature. The paper concludes with a lifecycle-oriented research agenda centered on bounded-autonomy specifications, adversarial trajectory generation, runtime monitoring, and audit-ready evidence structures. The central claim is that trustworthy deployment of agentic AI depends on validating trajectories in context rather than assessing isolated components alone.
发表机构
- University of Messina(墨西拿大学)
机构由 AI 辅助整理,请以论文原文为准。