不要凭直觉部署AI智能体:能力并非生产就绪
Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness
- The University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对AI智能体部署依赖能力而非生产就绪性的问题,提出ProofAgent Index(PAI)这一四维度治理就绪指标,经医疗、金融领域验证可区分风险,实现从直觉部署到可审计决策的转变。
AI中文摘要:
AI智能体正逐步进入生产工作流,它们检索信息、调用工具、维护状态并代表用户或组织执行操作,但许多发布决策仍依赖能力信号、演示或行为测试,这些无法体现智能体是否已准备好应对生产约束。因此,能力不等于生产就绪。本文提出了AI智能体的治理就绪指标——ProofAgent Index(PAI),PAI整合了部署证据的四个维度:评估、上下文、合规性和治理。评估衡量观测到的行为,上下文衡量塑造该行为的运行环境,合规性衡量与适用规则和控制措施的一致性,治理衡量组织是否能在运行期间授权、监控、审计和控制智能体。PAI在ProofAgent Harness内部实现,这是一个用于可审计AI智能体评估与治理的开源基础设施。在医疗和金融两个高度监管领域的验证表明,PAI具备保留的就绪信号,可区分高风险和低风险配置。结果显示,上下文工程会显著改变可靠性,能力可改善行为但不决定就绪性,治理证据必须保持可见而非被平均化。PAI将智能体发布从基于直觉的部署决策重构为可审计的就绪决策。
英文摘要:
AI agents are moving into production workflows where they retrieve information, call tools, maintain state, and act on behalf of users or organizations, but many release decisions still rely on capability signals, demos, or behavioral tests that do not show whether an agent is ready to operate under production constraints. Capability is therefore not production readiness. This paper introduces the ProofAgent Index (PAI), a governance readiness index for AI agents. PAI combines four dimensions of deployment evidence: Evaluation, Context, Compliance, and Governance. Evaluation measures observed behavior, Context measures the operating environment that shapes that behavior, Compliance measures alignment with applicable rules and controls, and Governance measures whether the organization can authorize, monitor, audit, and control the agent during operation. PAI is implemented inside ProofAgent Harness, an open source infrastructure for auditable AI agent evaluation and governance. Validation across two heavily regulated domains, healthcare and finance, shows that PAI carries held out readiness signal and separates higher risk from lower risk configurations. The results show that context engineering strongly changes reliability, capability improves behavior but does not determine readiness, and governance evidence must remain visible rather than averaged away. PAI reframes agent release from a faith based deployment decision into an auditable readiness decision.