发表机构
University of Trento; University of Turin; Politecnico di Torino(特伦托大学; 都灵大学; 都灵理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AgentLSD提出对抗性任务污染概念,通过CTF环境注入陷阱工件,评估六个AI安全智能体,发现陷阱显著增加回合与推理令牌,揭示干净性能低估脆弱性。
AI 中文摘要
用于安全的AI智能体会检查网页、源代码、日志、配置文件和命令输出。这些环境可能包含影响智能体行为的欺骗性工件。我们将此称为对抗性任务污染。提示注入依赖于攻击者提供的指令,而任务污染还包括非指令性证据,如虚假结果和诱饵端点。我们提出了AgentLSD,一个用于研究对抗性任务污染的受控框架。AgentLSD使用夺旗(CTF)挑战作为其实验环境。我们在保留预期CTF解决方案的同时,注入陷阱工件,如虚假标志、误导性提示、诱饵端点和隐藏线索。该框架支持配对的干净和陷阱增强实验,具有确定性陷阱生成、运行时注入、遥测和交付验证功能。我们在11个Web CTF挑战上评估了六个模型。在干净条件下,智能体捕获了41%的标志,且没有模型解决所有挑战。然后我们衡量了任务污染的影响。即使智能体仍然恢复标志,陷阱也会增加回合数(+20)和推理令牌数(+2k)。解决率的影响更加异质,因为某些模型-挑战对基本不受影响,而其他则跟随诱饵或提交错误的标志。这些结果表明,干净的CTF性能低估了对欺骗性任务证据的脆弱性。AgentLSD隔离了这种效应,并提供了一个可复现的基准来研究它。我们发布了框架、配置、陷阱规范和原始跟踪。
英文摘要
AI agents for security inspect web pages, source code, logs, configuration files, and command outputs. These environments may contain deceptive artifacts that influence the agent's behavior. We call this adversarial task contamination. Whereas prompt injection relies on attacker-supplied instructions, task contamination also includes non-instructional evidence, such as fake results and decoy endpoints. We present AgentLSD, a controlled framework for studying adversarial task contamination. AgentLSD uses Capture the Flag (CTF) challenges as its experimental environment. We inject trap artifacts, such as fake flags, misleading hints, decoy endpoints, and hidden cues, while preserving the intended CTF solution. The framework supports paired clean and trap-augmented experiments with deterministic trap generation, runtime injection, telemetry, and delivery verification. We evaluate six models on 11 web CTF challenges. In the clean condition, agents capture 41% of the flags, and no model solves every challenge. We then measure the impact of task contamination. Even when the agent still recovers the flag, traps increase the number of turns (+20) and reasoning tokens (+2k). Solve-rate effects are more heterogeneous, as some model-challenge pairs are largely unaffected while others follow decoys or submit wrong flags. These results show that clean CTF performance understates vulnerability to deceptive task evidence. AgentLSD isolates this effect and provides a reproducible benchmark for studying it. We release the framework, configurations, trap specifications, and raw traces.
CommentsTo be published in the 19th ACM Workshop on Artificial Intelligence and Security (AISec 2026) co-located with CCS 2026