发表机构
LinkedIn; Uber(领英; 优步)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出金融智能体可治理性基准Fiducia-bench,发现智能体分解会降低政策合规性,且弱化程度与模型能力相关,相关资源均开源。
AI 中文摘要
现有的智能体基准测试仅关注智能体是否完成了任务,而我们研究的是智能体是否在政策框架内完成任务。我们推出了Fiducia-bench,这是一个用于评估金融智能体可治理性的基准测试——即智能体是否在有义务时升级处理、在被要求时弃权(不执行),并留下可审计的追踪记录——并利用它研究了一个此前无基准测试涉及的问题:将智能体分解为组件是否会降低其治理水平?答案是肯定的,且机制具有特异性。一个组件发现的与政策相关的事实,在传递到需据此采取行动的组件时,会在边界处被弱化。在一项涉及100种KYC/AML任务变体、2种模型和3种架构的626轮实验中,32B开放权重模型在单循环基线设置下对已发现事实的弱化率为0%,在固定流水线设置下为56%,在编排器-子智能体架构下为85%(约束距离均为2)。更强的模型gpt-4.1-mini在相同条件下的弱化率为3%-6%,这表明分解带来的治理成本部分取决于模型能力。关键的是,同一机制会同时导致升级不足和过度升级,具体取决于被遗漏的事实是风险信号还是免责信号。该基准测试、所有任务及验证工具均为开源。
英文摘要
Existing agent benchmarks ask whether the agent finished the task. We ask whether it finished it within policy. We introduce Fiducia-bench, a benchmark for the governability of financial agents---whether they escalate when obligated, abstain when required, and leave an auditable trail---and use it to study a question no prior benchmark addresses: does decomposing an agent into components degrade its governance? It does, and the mechanism is specific. Policy-relevant facts discovered by one component are attenuated at the handoff boundary before reaching the component that must act on them. In a 626-episode experiment across 100 KYC/AML task variants, two models, and three architectures, a 32B open-weights model attenuated 0% of discovered facts under a single-loop baseline, 56% under a fixed pipeline, and 85% under an orchestrator-subagent architecture (all at constraint distance 2). A stronger model (gpt-4.1-mini) attenuated 3-6% under the same conditions, suggesting the governance cost of decomposition is partly a function of model capability. Critically, the same mechanism produces both under-escalation and over-escalation, depending on whether the dropped fact was a risk signal or an exculpating one. The benchmark, all tasks, and the verification harness are open-source
Comments8 pages, 3 tables, 1 figure. Preprint