arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体流程入侵检测:基于证据的运行时监控

Intrusion Detection for Agentic Processes: Evidence-Based Runtime Monitoring

Arslan Brömme

arXiv 2610.00151首次发表:更新:

AI 中文总结

本文提出智能体流程入侵检测系统(A-IDS),通过证据感知的运行时监控比较观察与期望基线,以检测智能体流程中的安全偏差,贡献为概念性框架。

AI 中文摘要

智能体部署日益将语言模型推理与检索、委派、工具执行、外部系统访问及人工审批相结合。因此,安全相关的偏差可能出现在一个不断演进的流程中,而非孤立的输入或动作。基于作者先前针对智能体流程的黑盒架构以及随后的证据-声明模型,本文提出了一种智能体流程入侵检测系统(A-IDS),这是一个具备证据感知能力的安全解释层,用于运行时入侵检测,其监控对象是智能体流程本身。A-IDS将证据支持的观察结果与受治理且版本化的期望基线进行比较,该基线涵盖工作流状态、授权、通信和强制事件。其概念贡献在于将动态到期的治理期望、与三值匹配分离的可见性、显式未解析的观察状态以及将证据状态与操作影响分离的有界发现相结合。该模型进一步将监控平面本身识别为攻击面,当对抗性内容通过看似合法的观察路径到达语义证据生产者时。某些观察结果可能在操作智能体自报告路径之外产生,但该模型不假设完全可观测性或普遍可信的捕获。提示注入既被视为输入安全问题,也被视为后续流程偏差和跨智能体影响路径的可能来源。A-IDS不从异常行为推断恶意意图,不将未观察事件视为未发生的证据,也不声称提出新的异常检测器、时序逻辑或溯源模型。其贡献是概念性的:它不验证实现、不展示实证检测性能、不建立因果归因,也不提供强制执行机制。

英文摘要

Agent deployments increasingly combine language-model inference with retrieval, delegation, tool execution, external-system access, and human approval. Security-relevant deviations can therefore emerge across an evolving process rather than in one isolated input or action. Building on the author's earlier black-box architecture for agentic processes and the subsequent evidence-claim model, this paper proposes an Agentic-Process Intrusion Detection System (A-IDS), an evidence-aware security interpretation layer for runtime intrusion detection whose monitored object is the agentic process itself. A-IDS compares evidence-supported observations with a governed and versioned expectation baseline for workflow state, authorization, communication, and mandatory events. Its conceptual contribution combines dynamically due governed expectations, visibility separated from three-valued matching, explicit unresolved observation states, and bounded findings that separate evidentiary status from operational impact. The model further identifies the monitoring plane itself as an attack surface when adversarial content reaches semantic evidence producers through otherwise legitimate observation paths. Some observations may be produced outside the operational agent's self-report path, but the model does not assume complete observability or universally trustworthy capture. Prompt injection is treated both as an input-security problem and as a possible origin of later process deviations and cross-agent influence paths. A-IDS does not infer malicious intent from anomalous behavior, does not treat an unobserved event as proof of non-occurrence, and does not claim a new anomaly detector, temporal logic, or provenance model. The contribution is conceptual: it does not validate an implementation, demonstrate empirical detection performance, establish causal attribution, or provide an enforcement mechanism.

Comments7 pages, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑