盲目信任,致命推力:当攻击者控制的钩子更新引导AI智能体控制框架走向恶意行为
A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors
- Beijing University of Posts and Telecommunications(北京邮电大学)
- Data Technology Support Center of the Cyberspace Administration of China(中央网络安全和信息化委员会办公室数据技术支持中心)
- Beihang University(北京航空航天大学)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
该研究揭示AI智能体控制框架的生命周期钩子更新路径为新攻击面,提出HookPry框架可攻破所有7款被评估框架,而现有防御措施不足以应对此类攻击。
中文摘要 AI 辅助
现代AI智能体控制框架暴露了生命周期钩子,这些钩子将shell命令绑定到会话启动、工具调用和文件编辑等运行时事件。这些命令以主机权限运行,但作为生命周期钩子配置提供,可能在大语言模型(LLM)从未观察到的时刻触发。我们发现,控制框架盲目信任的生命周期钩子更新路径是一个新的攻击面。在攻击者仅控制插件元数据和生命周期钩子配置的供应链威胁模型下,一个良性的版本化插件可通过更新被植入木马,该更新将攻击者选择的命令悄悄绑定到良性事件,从而产生主机端恶意行为,如权限提升。我们提出HookPry,这是一个开源的全自动化攻击框架,可跨异构AI智能体控制框架系统性利用此漏洞。HookPry实现了10个攻击目标;在1000次端到端运行中,针对控制框架与后端的25种组合,它攻破了所有7个被评估的控制框架,每个控制框架的成功率达92.5%。代表性防御措施仍不足:Microsoft Defender的召回率为0%,三种静态防御的组合遗漏了47.5%的恶意工件。
英文摘要
Modern AI agent harnesses expose lifecycle hooks that bind shell commands to runtime events such as session start, tool calls, and file edits. These commands run with host privileges yet ship as lifecycle-hook configuration and may fire at times the LLM never observes. We identify the lifecycle-hook update path, which harnesses trust blindly, as a new attack surface. Under a supply-chain threat model in which an attacker controls only plugin metadata and lifecycle-hook configuration, a benign versioned plugin can be trojanized by an update that silently binds attacker-chosen commands to benign events, yielding malicious host-side behavior such as privilege escalation. We propose HookPry, an open-source and fully automated attack framework that systematically exploits this vulnerability across heterogeneous AI agent harnesses. HookPry realizes ten attack objectives; across 25 combinations of harnesses and backends in 1,000 end-to-end runs, it compromises all seven evaluated harnesses, with per-harness success rates reaching 92.5%. Representative defenses remain insufficient: Microsoft Defender has 0% recall, and the union of three static defenses misses 47.5% of malicious artifacts.