arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ACEA:用于红队与蓝队大语言模型头对头测试的对抗性协同进化竞技场

ACEA: An Adversarial Co-Evolution Arena for Head-to-Head Red-Team and Blue-Team LLM Testing

Yi Ting Shen, Kentaroh Toyoda, Alex Leung

arXiv 2609.08256首次发表:更新:

发表机构

Vulcan Research, AIFT(AIFT火神研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ACEA平台,通过可插拔适配器连接红蓝队与目标LLM,利用LLM裁判评分,提供四组件支持对抗性评估,实现头对头攻击与防御的有效测试。

AI 中文摘要

针对大型语言模型(LLM)的自动化红队攻击和蓝队防御正在迅速发展。然而,攻击者和防御者是在隔离环境中构建和测试的,因此所得分数难以令人信服。为解决这一问题,我们提出了ACEA(对抗性协同进化竞技场),这是一个平台,它将可插拔的红队适配器和可插拔的蓝队适配器连接到共享的目标LLM,并通过LLM裁判对它们的攻击率和防御率进行评分。ACEA贡献了四个组件。首先,一个可插拔、与模型无关的竞技场。任何红队或蓝队项目都可以通过一个最小的HTTP协议(我们称之为ACEA标准适配器协议,ASAP)进行连接。该协议可以用任何语言编写,只要项目暴露该协议即可成为完全参与者。其次,一种为对抗性回合设计的评估方法。用规范秘密对目标进行种子设定,可以提供可验证的基准真相,从而区分真实泄漏与幻觉。我们还即使防御阻止了攻击,也会将每次攻击发送给目标,这可以独立于攻击是否被阻止来衡量攻击的原始威力。这些共同产生了每回合攻击强度和防御有效性的分解。第三,实时、游戏风格的视觉化,并附带详细的战斗结束报告,定位每个失败点。因此,评估成为改进红队或蓝队项目的可操作信号。第四,一个可选的上下文内改进循环,将每回合的结果转化为下一回合的建议提示。适配器可以跨回合适应而无需保持状态,只要它读取提示即可。我们描述了ACEA的设计以及红队和蓝队进行头对头评分的指标。

英文摘要

Automated red-team attacks and blue-team defenses for large language models (LLMs) are advancing quickly. However, attackers and defenders are built and tested in isolation, and the resulting scores are hard to trust. To tackle this, we present ACEA (Adversarial Co-Evolution Arena), a platform that connects a pluggable red-team adapter and a pluggable blue-team adapter to a shared target LLM and scores their attack and defense rates with an LLM judge. ACEA contributes four components. First, a pluggable, model-agnostic arena. Any red or blue project connects over a minimal HTTP protocol, which we call the ACEA Standard Adapter Protocol (ASAP). It can be written in any language, and a project that exposes nothing but the protocol is a full participant. Second, an evaluation methodology built for adversarial rounds. Seeding the target with canonical secrets gives verifiable ground truth that separates real leakage from hallucination. We also send each attack to the target even when the defense blocks it, which measures the attack's raw potency independently of whether it was stopped. Together these yield a per-round decomposition of attack strength and defense effectiveness. Third, a real-time, game-style visualization with a detailed end-of-battle report that localizes each failure. The evaluation thus becomes an actionable signal for improving a red or blue project. Fourth, an optional in-context improvement loop that turns each round's outcome into advisory hints for the next. An adapter can then adapt across rounds without keeping state, provided it reads the hints. We describe the design of ACEA and the metrics through which red and blue teams are scored head to head.

CommentsCode is available at https://github.com/VulcanLab/ACEA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑