Loopjacking:劫持人在环审批
Loopjacking: Hijacking Human-in-the-Loop Approval
浏览论文内容
中文总结 AI 辅助
本文提出 Loopjacking 攻击概念,指人类审批操作 A 但系统执行操作 B,通过复现多种智能体产品中的攻击,证明完整审批呈现和精确比较可防御此类攻击。
中文摘要 AI 辅助
人工审批通常被视为智能体执行关键操作前的最后一道安全边界。只有当提交审查的操作与后续被授权或释放的操作完全一致时,该边界才有意义。我们将这种绑定关系的失效称为 Loopjacking:人类批准了他们理解为操作 A 的内容,而实现却将该决策用于实质不同的操作 B。我们区分了两种变体。在基于表示的攻击中,B 在审批时已被编码但被省略或错误呈现;在审批后状态替换攻击中,人类看到正确的 A,而可变的工作流状态随后将其替换为 B。我们评估了一组有目的的已发布智能体产品。我们在七个测试的 Agno AgentOS 版本(截至 3.0.9)以及十二个测试的 LangGraph Agent Server 条件内存组合版本(截至 0.14.0)中复现了审批后替换攻击。我们在 OpenClaw 2026.2.23 中复现了表示不匹配,并在 2026.2.24 中验证了其被拒绝。OpenAI Agents SDK 0.22.0 和 0.22.2 提供了阴性对照:序列化延续保留了每次调用的精确绑定,并拒绝了被篡改的 B。这些结果并不估计生态系统的普遍性。它们表明,完整的规范化审批呈现和精确的使用时比较,或防止未授权的待处理状态变更,能够阻止所测试的攻击,同时保留合法执行。我们将此贡献与已有的关于误导性对话框、会话走私、动作绑定和授权连续性的工作区分开来。
英文摘要
Human approval is often treated as the last security boundary before an agent executes a consequential operation. That boundary is only meaningful if the operation presented for review is the operation later authorized or released. We call failures of this binding Loopjacking: a human approves what they understand as operation A, while the implementation uses that decision for a materially different operation B. We distinguish two variants. In a representation-based attack, B is already encoded but omitted or misrepresented at approval time; in a post-approval state-substitution attack, the human sees the correct A and mutable workflow state later replaces it with B. We evaluate a purposive set of released agent products. We reproduce post-approval substitution in seven tested Agno AgentOS releases ending at 3.0.9 and in 12 tested versions of a conditional in-memory LangGraph Agent Server composition ending at 0.14.0. We reproduce representation mismatch in OpenClaw 2026.2.23 and its rejection in 2026.2.24. OpenAI Agents SDK 0.22.0 and 0.22.2 provide a negative control: serialized continuation preserves exact per-call binding and rejects mutated B. These results do not estimate ecosystem prevalence. They show that complete canonical approval rendering and exact use-time comparison, or preventing unauthorized pending-state mutation, block the tested attacks while preserving legitimate execution. We separate this contribution from established work on misleading dialogs, session smuggling, action binding, and authorization continuity.