arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考人工智能系统的渗透测试:从资源妥协到行为目标违反

Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation

Mohammad Allahbakhsh, Mohammad Hassan Bahari, Moslem Attar-Raouf

arXiv 2607.14006首次发表:更新:

发表机构

Sensifai BV(Sensifai公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文重新思考人工智能系统渗透测试,将其定义为目标驱动的行为评估,涵盖多种对抗途径。提出测试工作流程,通过实例展示渗透可通过行为影响发生,为评估已部署人工智能系统中的对抗成功提供技术框架。

AI 中文摘要

传统的渗透测试评估对手是否能利用软件、基础设施、配置或操作控制中的弱点来实现与安全相关的妥协。这种范式对人工智能系统仍然必要,但已不充分。在这类系统中,对手可能影响提示、检索内容、传感器输入、训练数据、内存、工具或人机交互循环,以改变系统行为而不直接损害底层基础设施。本文将人工智能系统的渗透测试重新定义为目标驱动的行为评估。我们定义了人工智能系统及人工智能渗透。此定义保留了传统渗透测试,同时扩展到提示注入、间接提示注入、数据中毒、传感器操纵、检索中毒、工具滥用和代理失调等对抗途径。我们还提出了一个测试工作流程,包括识别操作目标、绘制人工智能控制行为、分析对抗影响面、定义行为失败标准、执行基于场景的测试以及报告将对抗行动与目标违反联系起来的证据。一个涉及人工智能安全运营中心助手的实例说明了渗透可能通过行为影响而非基础设施妥协发生。这些定义、工作流程和示例为评估已部署人工智能系统中的对抗成功提供了一个技术框架。

英文摘要

Penetration testing traditionally evaluates whether adversaries can exploit weaknesses in software, infrastructure, configurations, or operational controls to achieve security-relevant compromise. This paradigm remains necessary for AI-enabled systems, but it is no longer sufficient. In such systems, adversaries may influence prompts, retrieved content, sensor inputs, training data, memory, tools, or human-AI interaction loops to alter system behavior without directly compromising the underlying infrastructure. This paper reframes penetration testing for AI-enabled systems as objective-driven behavioral evaluation. We define an AI-enabled system as one in which learned models materially influence behavior affecting operational outcomes, and we define AI-enabled penetration as the feasible induction of AI-governed behavior that violates one or more operational objectives under an explicit threat model. This definition preserves conventional penetration testing while extending it to adversarial pathways such as prompt injection, indirect prompt injection, data poisoning, sensor manipulation, retrieval poisoning, tool misuse, and agentic misalignment. We further propose a testing workflow that identifies operational objectives, maps AI-governed behavior, analyzes adversarial influence surfaces, defines behavioral failure criteria, executes scenario-based tests, and reports evidence linking adversarial action to objective violation. A running example involving an AI-enabled security operations center assistant illustrates how penetration may occur through behavioral influence rather than infrastructure compromise. Together, the definitions, workflow, and example provide a technical framework for evaluating adversarial success in deployed AI-enabled systems.

Comments42 pages, 5 Tables, 21 references

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑