AgentS4D:基于大语言模型的工作空间智能体执行生命周期内运行时风险基准
AgentS4D: Benchmarking Runtime Risks across the Execution Lifecycle of LLM-Based Workspace Agents
浏览论文内容
中文总结 AI 辅助
研究推出AgentS4D基准,评估20种智能体配置在328个风险案例中的表现,发现任务完成不代表安全,需跨风险条件完整评估智能体配置。
中文摘要 AI 辅助
基于大语言模型(LLM)的工作空间智能体会在异构资源、外部工具和持久状态上执行有状态的多步骤工作流,因此必须在整个执行过程中从动作、副作用和状态变化等方面评估其安全性。尽管近期的基准在可执行安全测试和轨迹感知验证方面取得了进展,但它们很少统一说明风险从何处进入、如何引发不安全行为、针对哪些危害以及执行过程中支持证据出现在哪里。我们推出AgentS4D,这是一个用于全生命周期运行时安全评估的沙箱基准。其四维运行时安全框架采用6种风险引入源、6种诱导策略和9种目标危害来指导案例构建,同时7个生命周期检查点组织运行后证据。AgentS4D包含328个注入风险的案例。我们在这些案例上评估了4种 harness(Hermes、OpenClaw、Claude Code和Codex)与5种LLM后端(GPT-5.5、Gemini 3.1 Pro、DeepSeek-V4-Pro、MiniMax-M3和Qwen3.7-Plus)的全部20种组合,共产生6560次运行。总体而言,4461次运行(68.0%)触发了预设的不安全信号。在20种配置中,智能体系统的观测安全性随其harness-LLM配对以及风险引入方式而变化。当相同的诱导策略通过不同风险载体传递给智能体时,它们表现出明显不同的安全行为;当相同的目标危害通过不同载体和策略实现时,它们的响应也不同。此外,4344次运行(总体占比66.22%)属于不安全但已完成的情况。因此,任务完成不能等同于运行时安全,仅测试一种风险形式会掩盖重要弱点,评估应检查完整的智能体配置在不同风险条件下的表现并保留整个执行过程中的证据。
英文摘要
Large language model (LLM)-based workspace agents execute stateful, multi-step workflows across heterogeneous resources, external tools, and persistent state. Their safety must therefore be assessed from actions, side effects, and state changes throughout execution. Although recent benchmarks have advanced executable safety testing and trajectory-aware verification, they rarely provide a unified account of where risks enter, how they elicit unsafe behavior, which harms they target, and where supporting evidence appears during execution. We introduce AgentS4D, a sandboxed benchmark for lifecycle-wide runtime safety evaluation. Its four-dimensional runtime-safety framework uses six risk-entry sources, six induction strategies, and nine target harms to guide case construction, while seven lifecycle checkpoints organize post-run evidence. AgentS4D contains 328 risk-injected cases. We evaluate all 20 combinations of four harnesses (Hermes, OpenClaw, Claude Code, and Codex) and five LLM backends (GPT-5.5, Gemini 3.1 Pro, DeepSeek-V4-Pro, MiniMax-M3, and Qwen3.7-Plus) on these cases, yielding 6,560 runs. Overall, 4,461 runs (68.0%) trigger prespecified unsafe signals. Across the 20 configurations, the observed safety of an agent system varies with both its harness-LLM pairing and how risk is introduced. Agent systems exhibit markedly different safety behavior when the same induction strategy reaches them through different risk carriers. They also respond differently to the same target harm when it is realized through different carriers and strategies. Moreover, 4,344 runs (66.22% overall) are unsafe yet complete. Thus, task completion cannot establish runtime safety, and testing only one form of a risk can conceal important weaknesses. Evaluations should examine complete agent configurations across diverse risk conditions and retain evidence throughout execution.