arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过先发制硬化防止智能应用中的数据泄露

Data Leakage Prevention in Agentic Applications via Preemptive Hardening

Akansha Shukla, Emily Bellov, Parth Atulbhai Gandhi, Yuval Elovici, Asaf Shabtai

arXiv 2607.18847首次发表:更新:

发表机构

Faculty of Computer and Information Science; Ben-Gurion University of the Negev(计算机与信息科学系; 内盖夫本· Gurion大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究智能应用数据泄露问题,提出预部署管道,通过分析提示模板等识别泄露模式并生成补丁,经硬化和验证确保应用功能不受影响,评估显示能有效减少数据泄露。

AI 中文摘要

智能系统将基于大语言模型的规划与外部工具接口集成,可能因指令/数据边界故障和提示注入攻击导致数据泄露和工具滥用。在跨多个代码库和异构智能体的工作流程中持续实施所需控制具有挑战性。为此,我们提出了一个预部署管道,用于扫描、硬化和验证智能应用。该管道分析提示模板、工具接口和工具调用代码,识别导致泄露的模式并生成可操作的补丁。通过对抗性提示注入攻击和良性输入变化对硬化后的应用进行验证,确保缓解措施不会干扰预期行为。在硬化阶段,对高风险工具进行优先级排序并应用微创缓解措施,包括模式收紧、边界清理、基于允许列表的工具门控和最小权限检查。在验证阶段,管道自动生成模拟越狱、指令覆盖和工具针对性操纵的攻击输入以及良性任务变体,以确认硬化后应用的功能在修复后得以保留。我们在五个实际智能应用以及AgentDojo基准上对该管道进行了评估。结果表明,该管道能识别反复出现的导致泄露的模式并生成可集成的补丁,消除基本越狱和指令覆盖攻击导致的泄露,在压力诱导操纵条件下减少91%的泄露,且无需持续运行时策略执行。

英文摘要

Agentic systems integrate LLM driven planning with interfaces to external tools, making data leakage and tool misuse feasible via instruction/data boundary failures and prompt injection attacks. Enforcing required controls consistently is particularly challenging in workflows spanning many codebases and heterogeneous agents. To address this challenge in multi agentic systems, we present a pre-deployment pipeline for scanning, hardening, and validation of agentic applications. The pipeline analyzes prompt templates, tool interfaces, and tool-invocation code to identify leakage-enabling patterns and generate actionable patches. The hardened application is then validated through adversarial prompt injection attacks and benign input variations ensuring that mitigations do not disrupt intended behavior. In the hardening stage, high-risk tools are prioritized, and minimally invasive mitigations are applied, including schema tightening, boundary sanitization, allowlist-based tool gating, and least-privilege checks. In the validation stage, the pipeline automatically generates attack inputs that mimic jailbreaks, instruction overrides, and tool-targeted manipulation, along with benign task variants, to confirm that the functionality of the hardened application is preserved after remediation. We evaluated the pipeline on five real-world agentic applications, as well as on the AgentDojo benchmark. Across all applications, the proposed pipeline identified recurring leakage-enabling patterns and generated patches that can be integrated without disrupting the intended application behavior. The resulting modifications of application code were shown to eliminate leaks when targeted by basic jailbreak and instruction-override attacks, achieving a 100% reduction in leakage, and reduce leaks by 91% under conditions of stress-induced manipulation, without the need of continuous runtime policy enforcement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑