发表机构
Scale AI; University of California, Santa Cruz; Vanderbilt University Medical Center(Scale AI; 加利福尼亚大学圣克鲁兹分校; 范德堡大学医学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出READY框架,用于评估AI智能体是否适合企业工作流部署,通过临床审计案例研究揭示了仅自主性能差异无法体现的部署相关差异,为企业智能体部署决策提供依据。
AI 中文摘要
AI智能体在基准测试中表现良好,仍可能不适合部署。现有AI智能体基准测试衡量智能体能否完成现实专业工作,而企业部署关注的是不同问题:在可接受的人工监督下,智能体能否达到要求的可靠性水平,且成本在可容忍范围内。我们推出Reliable Enterprise Agent Deployment(READY,可靠企业智能体部署),这是一个用于评估AI智能体是否适合部署到企业工作流的框架。READY保留每个工作流自身对成功执行的定义,同时应用通用评估流程。给定一个智能体、一个工作流和一类候选监督策略,READY会衡量人机系统的可靠性和运营成本,选择满足指定可靠性目标的最低成本策略,并在保留的案例上进行统计评估,最终生成的部署配置文件描述了支持的操作点:可靠性、人工监督负担和成本。READY作为开放测试平台实现,将工作流规范、执行、评估和评估解耦,并在现有智能体评估基础设施上运行。在一项涵盖16个智能体系统和750个案例的端到端临床审计案例研究中,READY揭示了自主性能隐藏的差异:两个系统的自主准确率仅相差0.3个百分点(72.8% vs. 72.5%),但在评估的监督策略下,要达到相同的76%可靠性目标,分别需要39.2%和29.6%的人工审核。因此,READY将企业智能体评估从“智能体能多好地完成工作?”转变为“在什么条件下、以什么成本能可靠部署?”通过明确这些条件并使其可统计测试,READY为比较智能体系统、设定监督要求和做出基于证据的部署决策提供了基础。
英文摘要
An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, whereas enterprise deployment asks a different question: whether an agent can meet a required reliability level, under acceptable human oversight, and at tolerable cost. We introduce Reliable Enterprise Agent Deployment (READY), a framework for qualifying AI agents for deployment on enterprise workflows. READY preserves each workflow's own definition of successful execution while applying a common qualification procedure. Given an agent, a workflow, and a class of candidate oversight policies, READY measures the reliability and operating cost of the human-AI system, selects the minimum-cost policy that satisfies a specified reliability target, and statistically qualifies it on held-out cases. The resulting deployment profile characterizes the supported operating point: reliability, human-oversight burden, and cost. READY is implemented as an open testbed that decouples workflow specification, execution, evaluation, and qualification, and runs on existing agent-evaluation infrastructure. In an end-to-end clinical-audit case study spanning 16 agent systems and 750 cases, READY reveals differences hidden by autonomous performance: two systems separated by only 0.3 points in autonomous accuracy (72.8% vs. 72.5%) require 39.2% versus 29.6% human review, respectively, to qualify at the same 76% reliability target under the evaluated oversight policy. READY thus shifts enterprise agent evaluation from how well can the agent perform the work? to under what conditions, and at what cost, can it be reliably deployed? By making those conditions explicit and statistically testable, READY provides a basis for comparing agent systems, setting oversight requirements, and making evidence-based deployment decisions.