arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22712cs.AI

可信智能体人工智能:故障模式、缓解策略及自主LLM系统的生命周期框架

Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems

Fayeq Jeelani Syed, Rehan Ahmad, Ali Al Bataineh, Aakriti Adhikari

首次发表
浏览论文内容

中文总结 AI 辅助

本文从安全鲁棒性、对齐监督、透明度、隐私治理和合规五个维度,系统梳理了自主LLM智能体的故障模式与缓解策略,并提出涵盖规范、设计、训练、评估、部署和监控六阶段的可信智能体开发生命周期框架。

中文摘要 AI 辅助

基于大型语言模型构建的智能体人工智能系统能够进行多步骤规划、使用外部工具、在记忆中保留信息,并与其他智能体协调。这些能力使其比静态语言模型更有用,但也引入了新的安全和操作风险。来自网站、电子邮件、文档和数据库的不可信内容可能与系统指令进入同一上下文;持久记忆可能跨会话携带受损信息;而外部工具的访问可能将错误的模型响应转化为具有实际后果的现实世界行动。本文从五个相互关联的维度审视智能体人工智能的可信性:安全性与鲁棒性、对齐与人类监督、透明度与可审计性、隐私与数据治理,以及监管合规性。文章将关键故障模式(包括间接提示注入、后门触发器、目标泛化错误、记忆污染和跨会话数据泄露)组织成统一的分类体系。文章还考察了主要缓解方法,如指令层级、上下文隔离、焦点突出、基于过程的监督、受限工具使用和隐私保护记忆,同时区分了有实证支持的技术与仍主要停留在概念层面的技术。基于此分析,我们引入了可信智能体开发生命周期(TADL),这是一个涵盖规范、设计、训练、评估、部署和监控六个阶段的框架。对于每个阶段,TADL识别相关的信任活动、预期证据和基于风险的决策门。尽管TADL尚未经过实证验证,它为开发和评估更安全、更负责任的智能体系统提供了结构化基础。文章最后指出了当前基准测试中的空白,并概述了未来研究的优先事项。

英文摘要

Agentic AI systems built on large language models can plan over multiple steps, use external tools, retain information in memory, and coordinate with other agents. These capabilities make them more useful than static language models, but they also introduce new security and operational risks. Untrusted content from websites, emails, documents, and databases can enter the same context as system instructions; persistent memory can carry compromised information across sessions; and access to external tools can turn an incorrect model response into a consequential real-world action. This article reviews the trustworthiness of agentic AI across five interconnected dimensions: safety and robustness, alignment and human oversight, transparency and auditability, privacy and data governance, and regulatory compliance. It organizes key failure modes, including indirect prompt injection, backdoor triggers, goal misgeneralization, memory contamination, and cross-session data leakage, into a unified taxonomy. It also examines major mitigation approaches, such as instruction hierarchies, context isolation, spotlighting, process-based supervision, constrained tool use, and privacy-preserving memory, while distinguishing techniques supported by empirical evidence from those that remain largely conceptual. Building on this analysis, we introduce the Trustworthy Agent Development Lifecycle (TADL), a six-phase framework covering specification, design, training, evaluation, deployment, and monitoring. For each phase, TADL identifies relevant trust activities, expected evidence, and risk-based decision gates. Although TADL has not yet been empirically validated, it provides a structured foundation for developing and evaluating more secure and accountable agentic systems. The article concludes by identifying gaps in current benchmarks and outlining priorities for future research.

发表机构

  • Indiana University Indianapolis(印第安纳大学印第安纳波利斯分校)
  • Purdue University Northwest(普渡大学西北校区)
  • Higher Colleges of Technology(高等技术学院)
  • Yale University(耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

↑