治理边缘:通过混合本地-云多智能体框架实现商业财产与意外险承保自动化
Governing the Edge: Automating Commercial Property and Casualty Insurance Underwriting via a Hybrid Local-Cloud Multi-Agent Framework
AI总结:
针对商业财产与意外险承保中行政耗时问题,提出混合本地-云多智能体框架,在数据驻留约束下自动化流程,显著提速并保证合规与审计。
AI中文摘要:
商业财产与意外险(P&C)的承保人将30%至40%的时间花在行政工作上,而非风险判断上,且单次提交手工处理约需40分钟。我们提出了“治理边缘”(Governing the Edge),一个针对该层面的多智能体框架,其设计围绕数据驻留约束:敏感提交数据不得离开边界。十一个智能体和两个确定性控制节点(其中一个是人工升级中断)构成了一个跨两层的13节点LangGraph工作流。处理原始提交的智能体在本地运行于Gemma 2上,每个智能体通过记录每次调用的工具接口绑定(当前原型中使用合成桩);仅匿名化评分和非识别字段跨至云端的Claude Sonnet。合规规则被编码为图边上的条件,因此不合规的提交永远不会到达定价环节。我们端到端地走通了一个完整场景:一份商业汽车提交,其主要驾驶员有严重违规记录,展示了每个智能体调用、每次工具调用以及由此产生的路由决策。该框架在单个边缘设备上仅需几分钟即可处理原本40分钟的手工提交,主要耗时在于串行化的设备端推理。在硬性停止违规层级上,该框架正确且可重复地执行每条规则,因为硬性停止是对温度0下提取字段的确定性谓词;在20场景基准上的总体合规准确率为70%,其余差距集中在较软、基于判断的层级。它保留了完整的审计轨迹,并始终尊重隐私边界。我们发布了所有代码、合规规则、工具桩和合成数据集。
英文摘要:
Underwriters in commercial Property and Casualty (P&C) insurance spend 30 to 40% of their time on administrative work rather than risk judgment, and a single submission takes about 40 minutes by hand. We present Governing the Edge, a multi-agent framework for that layer, organized around a data-residency constraint: sensitive submission data must not leave the perimeter. Eleven agents and two deterministic control nodes, one a human-escalation interrupt, form a 13-node LangGraph workflow across two tiers. Agents touching raw submissions run locally on Gemma 2, each bound through a tool interface logging every invocation (synthetic stubs in the current prototype); only anonymized scores and non-identifying fields cross to Claude Sonnet in the cloud. Compliance rules are encoded as conditions on graph edges, so a non-compliant submission never reaches pricing. We walk through one full scenario end to end, a commercial auto submission whose principal driver carries serious violations, showing every agent call, every tool invocation, and the resulting routing decision. The framework processes a 40-minute manual submission in a few minutes on a single edge device, dominated by serialized on-device inference. On the hard-stop violation tier the framework enforces every rule correctly and reproducibly, since hard stops are deterministic predicates over fields extracted at temperature 0; overall compliance accuracy across the 20-scenario benchmark is 70%, with the remaining gap concentrated in softer, judgment-based tiers. It keeps a complete audit trail and respects the privacy boundary throughout. We release all code, compliance rules, tool stubs, and synthetic datasets.