arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FraudBench:针对自适应欺诈的基于政策的银行智能体压力测试

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

Dheeraj Mohandas Pai, Lu Xian

arXiv 2608.18136首次发表:更新:

AI 中文总结

研究人员推出基于τ²-bench框架和τ-Knowledge环境的可执行基准FraudBench,测试银行智能体应对自适应欺诈的安全性,评估发现4款智能体攻击安全性为49%-65%,资金骡和第一方欺诈是主要薄弱点。

AI 中文摘要

对话智能体如今可通过工具为终端用户执行操作,同时持有客户数据库和内部政策文件,呼叫方仅通过对话即可访问这些资源。银行业是最典型的案例:既能解答问题的同一智能体,也能修改联系信息、重置PIN码或转账,因此普通客户服务与授权、欺诈检测及政策合规性密不可分。现有的金融欺诈基准仅对静态交易或消息进行分类,通用智能体安全基准则针对提示注入或一般有害使用;没有任何基准测试基于政策的银行智能体在呼叫方通过对话操纵身份、授权和信任时,能否安全执行操作。我们推出FraudBench,这是一个可执行基准,基于τ²-bench双控制框架和τ-Knowledge银行环境构建。智能体和模拟呼叫方均通过工具在共享、可变的账户状态下运行,智能体可授予呼叫方对选定工具的访问权限;该环境提供698份内部政策文件构成的语料库,智能体必须从中检索所需内容。FraudBench包含150个预设对抗场景;所有报告的实验均使用冻结的公开集107个场景(其中90个属于10种欺诈机制,另有17个为链式自适应攻击),另有43个链式攻击作为保留项。安全性具有历史依赖性:单控制任务仅缺少一个前置条件,而自适应攻击会因早期的试探、许可或失败尝试,导致后续在本地有效的请求变得不安全。每个场景都标注了可观察证据、禁止操作、安全处置方式及干预点。对4个智能体在107个分级任务上的初步单轮评估显示,其攻击安全性介于49%至65%之间,其中资金骡欺诈和第一方欺诈是最常见的跨模型薄弱点。

英文摘要

Conversational agents now act for end users through tools while holding access to customer databases and internal policy documents that a caller can reach through dialogue alone. Banking is the clearest case: the same agent that answers a question can also change contact details, reset a PIN, or move money, so ordinary customer service is inseparable from authorization, fraud detection, and policy compliance. Existing financial-fraud benchmarks classify static transactions or messages, and general agent-safety benchmarks target prompt injection or generic harmful use; none test whether a policy-grounded banking agent safely acts when a caller manipulates identity, authorization, and trust over a conversation. We introduce FraudBench, an executable benchmark built on the $τ^2$-bench dual-control framework and the $τ$-Knowledge banking environment. Both the agent and the simulated caller act through tools over shared, mutable account state, and the agent may grant the caller access to selected tools; the environment exposes a 698-document internal policy corpus that the agent must retrieve from. FraudBench contains 150 authored adversarial scenarios; a frozen public set of 107 (90 across ten fraud mechanisms plus 17 chained adaptive attacks) is used for all reported runs, with 43 further chained attacks held out. Safety is history-dependent: single-control tasks satisfy every precondition but one, and adaptive attacks make a later, locally valid request unsafe because of an earlier probe, admission, or failed attempt. Each scenario is annotated with observable evidence, prohibited actions, safe dispositions, and intervention points. A preliminary single-trial evaluation of four agents on the 107 graded tasks yields attack-security between 49\% and 65\%, with money-mule and first-party fraud the most common cross-model weaknesses.

Comments9 Pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑