arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30441cs.CR

ECLIPSE:针对长视程智能体系统的自演化隐蔽提示注入攻击

ECLIPSE: Self-Evolving Stealthy Prompt Injection Attack against Long-Horizon Agentic Systems

Shiqian Zhao, Yangfan Zhou, Xinfeng Li, Runyi Hu, Yechao Zhang, Yi Xie, Tianwei Zhang, Luu Anh Tuan

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出ECLIPSE框架,结合隐蔽攻击轨迹合成与工具链引导,在LASE-Bench上实现高攻击成功率,现有安全措施难以可靠防御,凸显需开发更有效防御手段。

中文摘要 AI 辅助

近期,Codex、Claude Code、OpenClaw等大语言模型(LLM)智能体可通过重复工具调用规划并执行长视程任务,该能力也为提示注入创造了新机会。现有攻击要么将恶意目标置于单一明确指令中,易被检测;要么将意图分散到多个执行阶段,成功完成的可靠性低。本研究提出ECLIPSE,一种针对长视程智能体系统的自演化隐蔽提示注入框架,它结合直接用户提示注入与间接工具侧注入,包含两个组件:一是隐蔽攻击轨迹合成,利用沙箱生成并迭代验证候选工具链,再将验证后的链渲染为自然的单次提示作为直接指令;二是工具链引导,通过静态工作流编码(SWE)将该计划传递到目标环境,SWE在目标工具描述中嵌入状态转换线索,动态轨迹修正(DTC)则在执行偏离计划链时提供修正信号。为实现系统评估,我们进一步推出LASE-Bench,这是一个包含120个恶意任务和198种独特工具的长视程智能体安全基准,其中96.7%的任务至少需要5次工具调用。实验结果显示,ECLIPSE攻击效果极佳:无防御时攻击成功率高达96.7%,在常见安全过滤器下为69.2%,在防御设置中比最强基线高出27.5%;针对代表性防御措施的评估进一步表明,现有安全措施无法可靠防御该攻击,这凸显了开发更有效防御手段的必要性。

英文摘要

Recently, large language model (LLM) agents, such as Codex, Claude Code, and OpenClaw, have become capable of planning and executing long-horizon tasks through repeated tool calls. This capability also creates new opportunities for prompt injection. Existing attacks either place the malicious objective in one explicit instruction, making it easy to detect, or distribute the intent across multiple execution stages, making successful completion unreliable. In this work, we propose ECLIPSE, a self-evolving and stealthy prompt-injection framework for long-horizon agentic systems. ECLIPSE combines direct user-prompt injection with indirect tool-side injection through two components. On the one hand, Stealthy Attack Trajectory Synthesis uses a sandbox to generate and iteratively verify candidate tool chains, then renders a verified chain as a natural one-shot prompt to serve as the direct instruction. Then, Tool-Chain Steering transfers this plan to the target environment through Static Workflow Encoding (SWE), which embeds state-transition cues in target-tool descriptions, and Dynamic Trajectory Correction (DTC), which supplies corrective signals when execution deviates from the planned chain. To enable systematic evaluation, we further introduce LASE-Bench, a long-horizon agent-safety benchmark with 120 malicious tasks and 198 unique tools; 96.7% of its tasks make at least five tool calls. The experimental results show that ECLIPSE is highly effective: it achieves up to 96.7% attack success without defense and 69.2% under the common safety filter, exceeding the strongest baseline by 27.5% in the defended setting. Evaluations against representative defenses further show that existing safeguards do not reliably defend it, which raises the need for more effective defenses.

发表机构

  • Nanyang Technological University(南洋理工大学)
  • The Hong Kong Polytechnic University(香港理工大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑