arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DGF-Bench:用于模拟和审计针对多智能体治理委员会的欺骗行为的基准

DGF-Bench: A Benchmark for Simulating and Auditing Deception Against Multi-Agent Governance Boards

Jeremy Canale

arXiv 2609.34913首次发表:更新:

AI 中文总结

DGF-Bench提出一个多智能体治理委员会基准,通过注入欺骗性内容审计模型,发现任务对齐攻击可显著降低严格性,DGF分数评估防御能力。

AI 中文摘要

使用工具的语言模型智能体可以像治理委员会一样审查企业项目:它们阅读证据、应用书面规则并决定项目是否可以继续进行。部分证据来自与决策有利害关系的供应商和项目成员。DGF-Bench是一个基准,其中由智能体组成的委员会(专业门控和一个整合其决策的总门控)审查合成档案,同时攻击者将欺骗性内容植入组织不担保的证据中。档案根据61条可执行规则从规范事实生成,包含42条权威记录和32份叙述性文档;每个门控都被认证为可从这些记录中判定。攻击从不改变权威值,因此受攻击的档案保持其干净副本的参考决策。只有当智能体接收到注入并采取确切的注入行动(其在配对的干净档案上不会采取该行动)时,成功才可归因;DGF分数是模型阻止的适用固定攻击的比例。通过阅读文档和记录本身,六个模型中有五个在85个门控中的82到85个上是结果严格的(处置、发现、行动和授权全部正确)。在2,622次受攻击的门控运行中,七种直接命令、虚假数据和虚假权威攻击对这五个模型获得了一次可归因的成功,而模仿组织自身流程的任务对齐攻击通过了其中四个:引用虚假审查程序的记录注释将GPT-6 Luna Pro从34个结果严格门控降至6个,将DeepSeek V4 Pro从33个降至7个。DGF分数范围从96.2到26.9,而一个策略感知的自适应攻击者在记录中写入内容,成功对抗了六个模型中的五个。审批工具未执行任何伪造的审批,但被欺骗的智能体提交了规则禁止的审批。开源包dgf-bench通过一条命令计算DGF分数。

英文摘要

Tool-using language-model agents can review enterprise projects as governance boards do: they read the evidence, apply written rules and decide whether the project may proceed. Part of that evidence comes from suppliers and project members with a stake in the decision. DGF-Bench is a benchmark in which a board of agents (specialist gates and a General gate that consolidates their decisions) reviews synthetic dossiers while an attacker plants deceptive content in evidence the organization does not vouch for. Dossiers are generated from canonical facts under 61 executable rules, with 42 authoritative records and 32 narrative documents; every gate is certified decidable from those records. Attacks never change an authoritative value, so an attacked dossier keeps the reference decisions of its clean copy. A success is attributable only when the agent receives the injection and takes the exact injected action, which it does not take on the paired clean dossier; the DGF score is the share of applicable fixed attacks a model blocks. Reading documents and records themselves, five of six models were outcome-strict (disposition, findings, actions and authorization all correct) on 82 to 85 of 85 gates. Over 2,622 attacked gate runs, seven direct-order, false-data and false-authority attacks obtained one attributable success against these five, whereas task-aligned attacks imitating the organization's own process passed against four of them: a record note citing a fake review procedure lowered GPT-6 Luna Pro from 34 to 6 outcome-strict gates and DeepSeek V4 Pro from 33 to 7. DGF scores ranged from 96.2 to 26.9, and a policy-aware adaptive attacker writing in records succeeded against five of six models. The approval tool executed no forged approval, yet deceived agents submitted approvals that the rules forbid. The open-source package dgf-bench computes the DGF score with one command.

Comments52 pages, 7 figures, 15 tables. Project page: https://www.dgfbench.com/ ; code: https://github.com/jeremy1392/DGF-Bench ; package: https://pypi.org/project/dgf-bench/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑