arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

金融服务中的智能体可靠性档案

Agent Reliability Profiles in Financial Services

Mike Hsu, Medha Bankhwal, Béatrice Moissinac, Kevin Werbach, Lukasz Szpruch, Bennett Hillenbrand

arXiv 2610.04123首次发表:更新:

发表机构

MLCommons; Cambridge Judge Business School; Mezuro; Zendesk; University of Pennsylvania Wharton School; University of Edinburgh; The Alan Turing Institute(机器学习通用组织; 剑桥大学贾奇商学院; Mezuro公司; Zendesk公司; 宾夕法尼亚大学沃顿商学院; 爱丁堡大学; 艾伦·图灵研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出智能体可靠性档案,为金融服务中的智能体部署提供标准化保证证据,通过三级验证和基准测试提升可靠性评估与信任。

AI 中文摘要

AI智能体可以采取行动。有时,这些行动可能超出预期范围。智能体可靠性可以定义为确保智能体将保持在预期边界内并在限制范围内运行的保证。目前,在金融服务领域,尚无共享的框架或语言来描述、验证和基准测试智能体部署的可靠性。这使得金融机构、供应商和监管机构难以在规模化层面评估和信任智能体,从而限制了开发和采用的速度。智能体可靠性的标准化、共享表示将填补这一空白。本文介绍了智能体可靠性档案(Agent Reliability Profile),这是金融服务中智能体部署的每个智能体单位的保证证据。每个档案记录一个有界、可证伪的主张:该智能体系统在其运行边界内可靠地运行。我们将“运行边界”定义为智能体具有:(1)定义的自主层级,(2)定义的操作设计域,(3)定义的动作类别,以及(4)定义的控制包络。生产保证通过三个级别推进,而档案模式保持不变:档案构建器(Profile Builder)从机构证据中编译一级断言档案(Level 1 Asserted Profile),档案验证器(Profile Validator)在其自身环境中测试部署以生成二级验证档案(Level 2 Validated Profile),由合格独立评估者执行相同测试则生成三级验证档案(Level 3 Verified Profile)。另外,基准档案(Benchmarked Profile)报告在参考条件下跨机构可比较的结果。我们描述了架构、工件、保证阶梯、可比性标志、相关工具、评估方法、对金融机构和监管机构的应用、局限性以及分阶段实施计划。

英文摘要

AI agents can take actions. At times, those actions can go beyond what is intended. Agent reliability can be defined as assurance that an agent will stay within intended bounds and operate within limits. Today, there is no shared framework or language for describing, validating, and benchmarking the reliability of agentic deployments in financial services. This makes it difficult for financial institutions, vendors, and regulators to assess and trust agents at scale, thus limiting the pace of development and adoption. A standardized, shared representation of agent reliability would fill the gap. This paper introduces the Agent Reliability Profile, a per-agent unit of assurance evidence for agent deployments in financial services. Each Profile records a bounded, falsifiable claim, this agentic system reliably functions within its operating boundary. We define "operating boundary" as an agent having; (1) a defined autonomy tier, (2) a defined operational design domain, (3) defined classes of action, and (4) a defined control envelope. Production assurance progresses through three levels while the Profile schema remains constant: a Profile Builder compiles a Level 1 Asserted Profile from institutional evidence, a Profile Validator tests the deployment in its own environment to produce a Level 2 Validated Profile, and operation of the same tests by a qualified independent assessor produces a Level 3 Verified Profile. Separately a Benchmarked Profile reports results comparable across institutions under reference conditions. We describe the architecture, the artifact, the assurance ladder, the comparability flag, associated tools, an evaluation methodology, applications for financial institutions and supervisors, limitations, and a staged implementation program.

Comments35 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑