arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Agentic Commerce Bench:衡量可花钱智能体的欺诈检测

Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money

Ankit Srivastava, Debjyoti Paul

arXiv 2609.35886首次发表:更新:

发表机构

Gordon AI(戈登人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体自主花钱场景,提出欺诈分类体系、含二十类欺诈的ACB基准及开源检测堆栈gordonguard,实测中位付款0.007美元使人工审查成本达付款价值143倍。

AI 中文摘要

AI智能体现在拥有支出权限,无需每次操作的人工确认即可结算付款。由此产生的损失往往并非安全故障:拥有正确域名、正确结算地址并真实交付服务的交易对手可以收取超出应得金额的费用,而任何基于身份的检查都无法发现这一点。我们提出了三种用于衡量和减少该损失的工件。首先,一个智能体商务欺诈分类体系,将五个观察层级(智能体推理、电汇、结算轨道、交易对手、委托人)与每个层级可获得的请求级和历史级证据区分开来,并记录哪些层级可以观察到哪些攻击。其次,Agentic Commerce Bench(ACB),一个包含二十个欺诈类别的基准,这些类别由生产聚合数据、1,647个编目服务操作和1,068笔结算生成,其中六笔涉及完全声称身份的交易对手。第三,gordonguard,一个开源检测器堆栈和离线测试框架,操作员可借此审计智能体配置、无需账户即可重放恶意交易对手,并内联运行相同检测器。在干净的训练流量上根据规定的误报预算进行校准,得到6.5%的干净标记率,该结果在三个独立生成中复现,并使得二十个类别中的八个类别不优于随机猜测。在推理层可观察的四个类别上,一个广泛使用的智能体安全扫描器在其越狱检测面板上运行,四个类别均得分为零,而对作为对照提供的越狱样本正确评分为1.0。实测中位付款额为0.007美元,这对部署施加了硬性约束:一次人工审查的成本是其审查的付款价值的143倍。

英文摘要

AI agents now hold spend authority and settle payments without per-action human confirmation. The resulting loss is often not a security failure: a counterparty with the correct domain, the correct settlement address and a genuinely delivered service can charge more than it should, and no check keyed on identity will see it. We present three artefacts for measuring and reducing that loss. First, a taxonomy of agentic commerce fraud that separates five observation levels (agent reasoning, wire, settlement rail, counterparty, principal) from the request-level and history-level evidence available at each, and records which levels can observe which attacks. Second, Agentic Commerce Bench (ACB), a benchmark of twenty fraud classes generated from production aggregates, 1,647 catalogued service operations and 1,068 settlements, of which six involve a counterparty that is exactly who it claims to be. Third, gordonguard, an open-source detector stack and offline harness with which an operator can audit an agent configuration, replay hostile counterparties without an account, and run the same detectors inline. Calibrating to a stated false-positive budget on clean training traffic gives a 6.5% clean flag rate, replicated across three independent generations, and leaves eight of twenty classes no better than chance. On the four classes a reasoning layer can observe, a widely used agent security scanner run over its jailbreak-detection panel scores zero on all four, while correctly scoring 1.0 on a jailbreak supplied as a control. A measured median payment of $0.007 places a hard constraint on deployment: one human review costs 143 times the value of the payment it examines.

Comments12 pages, 4 figures. Code and data: https://github.com/BuildWithGordonAI/agentcommercebench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑