AI 中文总结
针对LLM编码智能体的PolicyGuard框架,以自然语言策略文件引导LLM分类用户提示,在冻结测试集和隐藏保留集上表现优异,泛化与跨模型能力突出。
AI 中文摘要
AI编码智能体接收自由格式的自然语言提示,这些提示可能无意中包含凭证、个人可识别信息(PII)或专有商业数据。现有的数据防泄漏(DLP)解决方案依赖僵化的正则表达式模式、模型微调或厂商管理的分类器,可定制性有限。我们提出PolicyGuard,这是一种预模型拦截框架,利用自然语言策略文件引导的LLM对用户提示进行分类。我们的核心贡献包括:(1)策略即提示范式,其中DLP分类标准完全由非工程师可编辑的明文策略文档定义,无需代码更改或模型重新训练;(2)具有模板家族级数据拆分、隐藏保留集和冻结测试集的密封评估协议,以严格评估泛化能力;(3)对2000个多语言提示的综合实证评估显示,在包含927个提示的冻结测试集上,有效拦截率(EBR)达96.5%,误报率(FPR)仅为3.0%,在包含217个提示的隐藏保留集上准确率达100%。信息匹配的基线实验表明,PolicyGuard的自然语言格式显著优于JSON格式的等效内容(McNemar卡方值=31.58,p<0.001),且大幅优于零样本分类(Cohen's h=0.915)。跨模型可移植性实验显示,同一策略无需修改即可在四种不同LLM上实现86.4%-96.5%的有效拦截率。
英文摘要
AI coding agents accept free-form natural language prompts that may inadvertently contain credentials, personally identifiable information (PII), or proprietary business data. Existing data loss prevention (DLP) solutions rely on rigid regex patterns, model fine-tuning, or vendor-managed classifiers with limited customizability. We present PolicyGuard, a pre-model interception framework that classifies user prompts using an LLM guided by a natural language policy file. Our key contributions are: (1) the policy-as-prompt paradigm, where DLP classification criteria are defined entirely in a plaintext policy document editable by non-engineers without code changes or model retraining; (2) a sealed evaluation protocol with template-family-level data splits, hidden holdouts, and frozen test sets to rigorously assess generalization; and (3) a comprehensive empirical evaluation across 2,000 multilingual prompts demonstrating 96.5% effective block rate (EBR) with only 3.0% false positive rate (FPR) on a frozen test set of 927 prompts, and perfect 100% accuracy on a 217-prompt hidden holdout. Information-matched baseline experiments show that PolicyGuard's natural language format significantly outperforms equivalent content in JSON format (McNemar chi-squared = 31.58, p < 0.001) and dramatically outperforms zero-shot classification (Cohen's h = 0.915). Cross-model portability experiments demonstrate that the same policy achieves 86.4-96.5% EBR across four different LLMs without modification.
CommentsThis paper is being withdrawn because it was submitted prior to completion of a required institutional review process. The authors intend to resubmit after the review is complete