arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为关键系统设计可靠的智能体人工智能

Engineering Trustworthy Agentic AI for Critical Systems

Omar Al-Refai, Ibrahim Shahbaz, Adam Ali Husseinat, Michael Mandulak, Jaewon Kim, Eman Hammad

arXiv 2607.18548首次发表:更新:

发表机构

Texas A&M University; Innovations in Systems Trust and Resilience (iSTAR) Laboratory; Texas A&M Institute of Data Science - Security, Privacy and Resilience for Trusted AI (SPARTA) Thematic Lab.; Texas A&M Global Cyber Research Institute (GCRI)(德克萨斯A&M大学; 系统信任与弹性创新(iSTAR)实验室; 德克萨斯A&M大学数据科学研究所 - 可信人工智能的安全、隐私与弹性(SPARTA)主题实验室; 德克萨斯A&M大学全球网络研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对关键工程领域智能体人工智能,将可靠性作为首要工程属性,采用围绕五个维度的可靠性模型及保证工作流程,综述相关架构、机制等,经多领域检验,指出其可靠性是单一问题并勾勒跨域保证框架路径。

AI 中文摘要

能够自主感知、规划、使用工具和进行多步行动的智能体人工智能系统,越来越多地被应用于关键工程领域,这些领域的决策会带来物理、操作或经济后果。本综述弥补了当前文献中的一个空白,将可靠性(即在工程实践实际要求的约束下,智能体行为能否被验证、审核和信任)作为首要工程属性,而非仅通过任务能力来评估智能体人工智能。研究采用围绕五个交叉维度组织的可靠性模型:安全与约束满足;稳健性与可靠性;透明度与可解释性;问责与可审计性;隐私与安全性。这被映射到一个从感知到审核的智能体保证工作流程中。在此基础上,对智能体系统架构、威胁、具体信任机制和定量指标进行了综述,以便直接应用于智能体系统的开发和评估。然后在四个受约束的工程领域:电力系统、自动驾驶车辆/机器人/无人机、高性能计算和通信网络中检验这些原则,识别出反复出现的设计模式、共享的故障模式和特定领域的差距。综合这些领域的情况,表明智能体人工智能的可靠性是一个单一问题,并勾勒出一条通向类似于成熟安全关键工程领域使用的分级认证制度的可重用跨域保证框架的路径。

英文摘要

Agentic artificial intelligence systems, capable of autonomous perception, planning, tool use, and multi-step action, are increasingly proposed for critical engineering domains where decisions carry physical, operational, or economic consequences. This survey addresses a gap in current literature by treating trustworthiness, whether agentic behavior can be verified, audited, and trusted under the constraints that engineering practice actually requires, as a first-class engineering property, rather than evaluating agentic AI by task capability alone. The study adopts a trustworthiness model organized around five cross-cutting dimensions: safety and constraint satisfaction; robustness and reliability; transparency and interpretability; accountability and auditability; and privacy and security. This is mapped onto an agentic assurance workflow spanning perception through audit. Building on this foundation, agentic systems architectures, threats, concrete trust mechanisms, and quantitative metrics are surveyed for direct application in agentic systems development and evaluation. These principles are then examined across four constraint-bound engineering domains: power systems, autonomous vehicles/robotics/UAVs, high-performance computing, and communication networks, identifying recurring design patterns, shared failure modes, and domain-specific gaps. Synthesizing across those domains, agentic AI trustworthiness is shown to be a single problem, with a path outlined toward a reusable, cross-domain assurance framework analogous to the graded certification regimes used by mature safety-critical engineering fields.

CommentsThis work has been submitted to the IEEE for possible publication

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑