发表机构
School of Computing, Wichita State University; Department of Computer Science, American International University-Bangladesh(威斯康星州立大学计算机学院; 孟加拉国国际大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究网络操作中工具使用大型语言模型代理的间接提示注入问题,提出NetInjectBench基准测试。通过多种方法评估,发现单纯执行不安全率高,不同方法可降低该率,元数据感知策略门效果好,表明网络操作代理需执行时授权边界。
AI 中文摘要
使用工具的大型语言模型(LLM)代理对网络操作很有吸引力,但工单、警报、日志、运行手册和ChatOps消息可能携带间接提示注入。我们提出了NetInjectBench,这是一个包含130个场景的基准测试,它为网络操作工具使用分离了不可信工件文本、可信策略元数据和评估标签。样本包含40个良性、40个弱攻击、40个强攻击和10个批准的高影响变更场景;每个场景都使用Qwen2.5 - 7B、Llama3.1 - 8B和Mistral - 7B进行评估。在240个攻击实例中,单纯执行的不安全工具操作率达到82.50%。仅提示安全、自我提醒、突出显示和两遍LLM判断分别将该率降至25.63%、21.67%、18.33%和10.00%。静态允许列表达到5.00%,但阻止了所有批准的更改,在批准的情况下有用性为0.00%,过度阻止率为100.00%。在规定的元数据完整性假设下,元数据感知策略门产生了0/240个不安全攻击操作,95%威尔逊上限为1.58%,同时保留了99.17%的攻击场景有用性和100.00%的批准更改有用性。研究结果表明,网络操作代理除了提示级指令卫生外,还需要执行时授权边界。
英文摘要
Tool-using large language model (LLM) agents are attractive for network operations, but tickets, alerts, logs, runbooks, and ChatOps messages can carry indirect prompt injections. We present NetInjectBench, a 130-scenario benchmark that separates untrusted artifact text, trusted policy metadata, and evaluation labels for network-operation tool use. The sample contains 40 benign, 40 weak-attack, 40 strong-attack, and 10 approved high-impact change scenarios; each is evaluated with Qwen2.5-7B, Llama3.1-8B, and Mistral-7B. Across 240 attack instances, naive execution reached an 82.50% unsafe tool-action rate. Prompt-only safety, Self-Reminder, Spotlighting, and a Two-Pass LLM Judge reduced this rate to 25.63%, 21.67%, 18.33%, and 10.00%, respectively. Static allowlisting reached 5.00% but blocked all approved changes, yielding 0.00% usefulness and 100.00% overblocking on approved cases. Under the stated metadata-integrity assumption, the metadata-aware policy gate produced 0/240 unsafe attack actions, with a 95% Wilson upper bound of 1.58%, while preserving 99.17% attack-scenario usefulness and 100.00% approved-change usefulness. The findings show that network-operation agents need execution-time authorization boundaries alongside prompt-level instruction hygiene.