RedEvoAgent:具备经验驱动技能演化的自动红队代理
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
- City University of Hong Kong(香港城市大学)
- Shenzhen MSU-BIT University(深圳北理莫斯科大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对LLM代理的越狱攻击风险,提出RedEvoAgent黑盒红队代理,通过提炼攻击轨迹实现技能演化,在多基准实验中性能优于基线,工具效率更高且可迁移。
中文摘要 AI 辅助
基于大语言模型(LLM)的代理正越来越多地被部署到产品级执行框架中,其越狱攻击会触发有害工具使用及持久状态变更,带来的风险远超单纯的不安全文本生成。现有自动红队方法通常依赖固定攻击,而近期的代理攻击者会协调多种越狱工具,通过基于轨迹的检索展现出更强潜力。但这类检索会因检索偏差和工具归因不明确而重复使用误导性经验,且完整轨迹会增加上下文开销并降低可解释性。我们提出RedEvoAgent,这是一种黑盒红队代理,它将跨案例攻击轨迹提炼为简洁、人类可读的攻击技能。该攻击技能通过工具有效性分析和用于技能更新的工具归因决策,以及仅保留提升验证性能的更新的验证棘轮机制实现自适应演化。在多个基准、目标模型和目标执行框架上的实验表明,RedEvoAgent的性能优于固定攻击和代理基线,提升了工具效率,且可在攻击者模型和目标执行框架间迁移。
英文摘要
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-teaming methods often rely on fixed attacks, while recent agentic attackers coordinate multiple jailbreak tools and show stronger potential through trajectory-based retrieval. However, such retrieval can reuse misleading experiences due to retrieval bias and unclear tool credit, and full trajectories add context overhead while reducing interpretability. We propose RedEvoAgent, a black-box red-teaming agent that distills cross-case attack trajectories into a concise, human-readable attack skill. The attack skill adaptively evolves through tool-effectiveness profiling and Deciding-Tool Attribution for skill updates, and a validation ratchet that retains only updates improving validation performance. Experiments on multiple benchmarks, target models, and target execution harnesses show that RedEvoAgent outperforms fixed and agentic baselines, improves tool efficiency, and transfers across attacker models and target execution harnesses.