arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于人工智能安全评估的对抗性提示框架

Adversarial Prompting Framework for AI Safety Assessment

Yash Bhatnagar, Kunal Banerjee, Anirban Chatterjee

arXiv 2607.13453首次发表:更新:

发表机构

Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对人工智能尤其是生成式人工智能应用增加带来的安全问题,提出对抗性提示框架,通过生成多复杂程度的对抗性提示评估模型弹性,在企业环境中实现自动化测试并获定量指标及差异结果。

AI 中文摘要

近年来,人工智能(AI)尤其是生成式人工智能(GenAI)在各行业的应用显著增加。然而,这些模型的使用也可能使系统面临不同恶意行为者的新型网络攻击,对抗性提示攻击(APA)就是此类威胁中最突出的例子之一。本文提出了一个对抗性提示框架(APF)来全面评估人工智能安全。该框架通过生成多个复杂程度的结构化对抗性提示,从直接有害请求到基于高级编码的攻击,系统地评估人工智能模型的弹性。我们的实现展示了这种方法在企业环境中的实际应用,提供了具有定量安全评估指标的自动化测试能力。结果表明,不同攻击向量下模型漏洞存在显著差异,编码提示在绕过安全机制方面成功率最高。

英文摘要

Artificial Intelligence (AI), especially Generative AI (GenAI), adoption has increased in industries significantly in recent years. However, the use of these models may also expose systems to new forms of cyberattacks by different malicious actors -- adversarial prompt attack (APA) being one of the most prominent examples of such threats. This paper presents the implementation of an Adversarial Prompting Framework (APF) for a comprehensive assessment of AI safety. The framework systematically evaluates the resilience of the AI model through the generation of structured adversarial prompts at multiple sophistication levels, from direct harmful requests to advanced encoding-based attacks. Our implementation demonstrates the practical application of this methodology in enterprise environments, providing automated testing capabilities with quantitative security assessment metrics. The results indicate significant variations in the model vulnerabilities across different attack vectors, with encoded prompts presenting the highest success rates in bypassing safety mechanisms.

Comments3 pages, 1 figure, presented as a poster at International Conference on Data Science (CODS), December 17-20, 2025, Pune, India

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑