arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24662cs.AI

DUMA-Bench:用于评估LLM智能体安全性的双控制多智能体基准

DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security

Ivan Aleksandrov, German Kochnev, Sabrina Sadiekh, Yaroslav Rogoza

首次发表
浏览论文内容

中文总结 AI 辅助

DUMA-Bench提出双控制交互基准,扩展tau2-bench覆盖八类漏洞,评估14个模型,发现双控制使攻击成功率从26.9%升至41.1%,揭示安全性源于模型、用户与环境交互。

中文摘要 AI 辅助

基于LLM的智能体越来越多地在与用户、工具和外部系统交互的环境中运行。然而,大多数安全评估假设用户是被动的且控制是静态的,忽略了塑造真实智能体行为的交互动态。我们引入了\textbf{DUMA-Bench},一个用于在\emph{双控制}交互下衡量智能体安全性的基准和评估协议,其中智能体和用户都能影响共享的环境状态。DUMA-Bench扩展了$\tau^2$-bench~\cite{barres2025tau},增加了覆盖八类漏洞类别的对抗性环境,包括RAG投毒、跨智能体操纵和不安全输出处理。我们评估了来自五个模型家族(OpenAI、Anthropic、DeepSeek、Qwen和this http URL)的\textbf{14个模型},涵盖八个领域和多种用户行为模式。在我们的实验中,引入双控制交互将攻击成功率从\textbf{26.9\\%}提高到\textbf{41.1\\%}。这些结果表明,智能体安全性不仅仅是模型的一种属性,而是从模型、用户和环境之间的交互中涌现出来的。DUMA-Bench为研究真实智能体部署中的安全性提供了一个缺失的评估层。

英文摘要

LLM-based agents increasingly operate in environments where they interact with users, tools, and external systems. Yet most security evaluations assume passive users and static control, ignoring the interactive dynamics that shape real agent behavior. We introduce \textbf{DUMA-Bench}, a benchmark and evaluation protocol for measuring agent security under \emph{dual-control} interaction, where both the agent and the user can influence the shared environment state. DUMA-Bench extends $τ^2$-bench ~\cite{barres2025tau} with adversarial environments covering eight vulnerability classes, including RAG poisoning, cross-agent manipulation, and unsafe output handling. We evaluate \textbf{14 models from five model families} (OpenAI, Anthropic, DeepSeek, Qwen, and Z.ai) across eight domains and multiple user-behavior regimes. Across our experiments, introducing dual-control interaction increases the attack success rate from \textbf{26.9\%} to \textbf{41.1\%}. These results show that agent security is not solely a property of the model but emerges from the interaction between the model, the user, and the environment. DUMA-Bench provides a missing evaluation layer for studying security in realistic agent deployments.

发表机构

  • ITMO University(ITMO大学)
  • Hive Trace Lab(Hive Trace实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑