arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28666cs.CLcs.LG

核查问题:受监管企业在AI产品交付前必须满足什么条件

The Checking Problem: What must be true before AI ships in a regulated firm

Prerit Ahuja

首次发表
浏览论文内容

中文总结 AI 辅助

本文研究受监管金融服务领域AI项目停滞的机制,通过多模型多配置实验发现,AI工作流价值更多取决于人类审核量,而非正确率,且审核负担可通过样本外估算测量。

中文摘要 AI 辅助

企业AI项目停滞的比率被广泛提及却缺乏合理解释,本文对该机制进行了研究。在受监管的金融服务领域,6种日常执行的文档密集型工作流被应用于4类模型家族和3种工具配置,每种配置重复执行3次,共产生72种配置下的5093个已评分输出元素。每种配置接受两次评估:一次针对演示基准(即单个案例的单次正确运行),另一次针对生产基准(要求持续准确性、重复运行的可复现性、可验证归因及携带信息的置信信号)。72种配置中,57种通过了演示基准,32种通过了生产基准,存活率为56.1%。随后,本文计算了每种配置带来的审核负担,该负担是通过样本外估算而非事后回溯得出的:不输出置信度的工具需要对100%的输出进行审核,因为它无法为审核人员提供分类依据;要求工具引用来源并说明置信度,可将该比例降至49%,同时在20种配置中的17种内保持剩余误差容限;增加自验证环节会使延迟达到普通配置的2.3倍,通过率为44%,且是唯一无法保持误差容限的配置。实际意义在于,AI工作流的价值较少取决于其正确的频率,而更多取决于人类仍需检查的工作量,且这一特性可测量却很少被测量。

英文摘要

Enterprise AI programmes stall at a rate that is widely quoted and poorly explained. This paper measures the mechanism. Six document-heavy workflows of the kind performed daily in regulated financial services were run across four model families and three tool configurations, three times each, producing 5,093 scored output elements across 72 configurations. Each configuration was assessed twice: against a demonstration bar, being a single correct run on a single case, and against a production bar requiring sustained accuracy, reproducibility across repeats, verifiable attribution, and a confidence signal that carries information. 57 of 72 configurations cleared the demonstration bar and 32 cleared the production bar, a survival rate of 56.1%. The paper then computes the review burden each configuration imposes, estimated out of sample rather than with hindsight. A tool that states no confidence requires review of 100% of its output, because it offers a reviewer no basis for triage. Requiring the tool to cite its sources and state a confidence reduces that to 49% while holding the residual error tolerance in 17 of 20 configurations. Adding a self-verification pass costs 2.3 times the latency of the plain configuration, reaches 44%, and is the only configuration that fails to hold the error tolerance. The practical implication is that the value of an AI workflow is set less by how often it is right than by how much of it a human must still check, and that the second property is measurable and rarely measured.

补充信息

↑