发表机构
Baylor University(贝勒大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对金融科技中智能体AI治理研究不足的问题,构建多层次治理理论,通过三项研究验证其机制,发现可验证性缺口是核心约束,该框架可扩展至其他高风险领域。
AI 中文摘要
金融机构正将重要决策委托给智能体人工智能系统,这些系统可分解目标、协调模型与工具,且几乎无需监督即可执行。然而,金融科技领域的智能体人工智能治理研究不足。我们认为,具有约束力的治理约束并非能力,而是可验证性。我们将可验证性缺口定义为:验证委托权限所需的内容,与决策后保留的可解释性和可复现性之间的差距。该缺口与验证者、证据标准及审计滞后时间相关。我们针对智能体人工智能构建了多层次治理理论,并在三项研究中测试其机制,涉及9个模型版本,从30亿参数的本地模型到商业前沿系统。研究1显示,提供商发布的版本会改变历史金融行为,控制回放需求属于提供商:前沿模型完全拒绝温度(temperature)、核采样(top_p)和top_k参数,且不暴露随机种子。在每个端点允许的最严格控制下,本地模型复现了320次执行中的320次,托管模型复现了320次中的319次及960次中的959次。研究2显示,编排是潜在的策略层,架构会改变最终行为,且任何配置、任何规模下都没有重复的执行记录。前沿模型复现自身行为的频率高于本地模型,但其记录并不更好,且损失了相当比例的区分度。能力带来更高的起点,而非可审计性。研究3显示,两个确定性信用模型版本各自完美复现当前行为,但当前版本无法恢复历史行为。我们将可复现性概念化为治理概况,而非标量,从而产生依赖证据的委托:仅在保留的证据能证明权限行使合理时,该权限才具备可辩护性。除金融领域外,该框架还可扩展至其他需要可审计性的高风险领域。
英文摘要
Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top_p and top_k outright and exposes no random seed. Under the tightest controls each endpoint allows, a local model reproduced 320 of 320 executions, hosted models 319 of 320 and 959 of 960. Study 2 shows that orchestration is a latent policy layer. Architecture changes final actions, and no execution record repeated in any configuration at any scale. The frontier model reproduces its own actions more often than the local ones, its record no better, and loses a comparable share of its differentiation. Capability buys a higher starting point, not auditability. Study 3 shows two deterministic credit-model versions each reproduce their current action perfectly, yet the current cannot recover a historical one. We conceptualize reproducibility as a governance profile, not a scalar, yielding evidence-contingent delegation: authority is defensible only while retained evidence substantiates its exercise. Beyond finance, the framework extends to other high-stakes domains requiring auditability.
Comments40 pages, 15 figures