arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CASE框架:用于管控企业智能体AI的多学科控制架构

The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI

Srinivas Telukunta, Georgios Nektarios Lilis, Lucio Baron

arXiv 2608.10153首次发表:更新:

发表机构

Cornell University; Johns Hopkins University; AI71(康奈尔大学; 约翰斯·霍普金斯大学; AI71)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对企业智能体AI管控的“涌现缺口”问题,提出含四层架构的CASE框架,经实证验证可提升治理有效性,满足欧盟AI法案要求。

AI 中文摘要

企业部署自主AI智能体的速度远超其管控能力,现有方法将单一学科(通常是为确定性自动化构建的DevSecOps)延伸至所有智能体规模。本文认为,智能体AI治理是四个问题而非一个,每个问题都有成熟的管控科学:CASE框架将控制理论应用于单个智能体(意图为设定值、护栏为反馈、评估为观测),将复杂自适应系统理论应用于智能体集合(其中涌现性使单智能体保证非组合性),将监督控制论应用于人机团队(其中必要多样性定律表明,无辅助的人工监督在结构上失效),将工程运营应用于智能体舰队(将错误预算延伸至决策质量,使自主性成为受控变量)。本文对各层进行形式化,推导跨层耦合条件,包括零接触部署悖论——某一层的卓越表现会对其他层造成压力,并追溯20余种企业管控措施至其经典构造。三项实证研究验证了该论点:82%的已记录生产智能体故障为多层轨迹;22种生态系统工具中无一种提供完整的第2层(涌现性)覆盖;所有35个评分的公开部署均处于最低成熟度区间。本文将这种不匹配(涌现性层的风险已实现,而能力几乎未提供且实践缺失)命名为“涌现缺口”。一个五级成熟度模型(含非补偿性瓶颈加权指数及评估工具)将CASE转化为科学而非流程成熟度模型,基于生产企业智能体平台。随着《欧盟AI法案》第14条将有效人工监督定为法律要求,只有满足必要多样性的架构才能使监督成为现实而非形式主义。

英文摘要

Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typically DevSecOps built for deterministic automation, across every scale of agency. We argue that agentic AI governance is four problems, not one, each with a mature governing science. The CASE framework assigns Control theory to the individual agent (intent as setpoint, guardrails as feedback, evaluation as observation), complex Adaptive systems theory to agent collectives (where emergence makes single-agent assurance non-compositional), Supervisory cybernetics to human-agent teams (where the Law of Requisite Variety shows unaided human oversight fails structurally), and Engineering operations to fleets (extending error budgets to decision quality so autonomy becomes a controlled variable). We formalize each layer, derive cross-layer coupling conditions, including a zero-touch deployment paradox where excellence at one-layer strains the others, and trace twenty-plus enterprise controls to their classical constructs. Three empirical studies validate the thesis: 82 percent of documented production agent failures are multi-layer trajectories; none of 22 ecosystem tools offers full Layer 2 (emergence) coverage; and all 35 scored public deployments fall in the lowest maturity band. We name this mismatch, risk realized at the emergence layer against capability barely offered and practice absent, the Emergence Gap. A five-level maturity model with a non-compensatory bottleneck-weighted index and assessment instrument operationalizes CASE as a scientific rather than process maturity model, grounded in production enterprise agentic platforms. As EU AI Act Article 14 makes effective human oversight a legal requirement, only architectures satisfying requisite variety can make oversight real rather than ceremonial.

Comments34 pages in total, 4 figures, 2 tables, code at https://github.com/srinivastelukunta/case_framework_arxiv_codes

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑