arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缺失的层次:人工智能监督的规范基础设施

The Missing Layer: Specification Infrastructure for AI Oversight

Satyam Kumar, Saurabh Jha

arXiv 2607.24866首次发表:更新:

AI 中文总结

研究指出人工智能安全监督存在协调差距,提出两轴分类法。针对规范层缺乏成熟学科标志的问题,给出六项设计原则,以CARMA为例展示如何让规范可组合,助独立团队构建缺失部分实现有效监督。

AI 中文摘要

人工智能安全存在缺失层次。可解释性、形式方法、安全工程、评估方法和强化学习安全都有大量工作成果,但这些工件无法组合成可部署的监督机制。我们将此诊断为协调差距而非研究差距,并提出两轴分类法。第二层(规范)缺乏成熟工程学科的四个标志,我们提出六项设计原则,并通过具体示例和参考架构进行说明。现有系统各处理第二层的一部分,将它们视为共享层的片段可便于组合。我们引入CARMA作为自主ETL代理的第二层原型来证明。

英文摘要

AI safety has a missing layer. Interpretability, formal methods, security engineering, evaluation methodology, and reinforcement-learning safety each produce substantial work, but the resulting artifacts do not compose into deployable oversight: every team fielding an agentic system builds its own audit schema, policy dialect, monitoring stack, and escalation path, mostly reinventions of patterns understood elsewhere. We diagnose this as a coordination gap, not a research gap, and propose a two-axis taxonomy: five technical layers (Legibility, Specification, Mediation, Evaluation, Escalation) crossed with six concerns spanning alignment, robustness, adversarial defense, security, governance, and accountability, populating the resulting 5x6 matrix with existing work. Layer 2 (Specification), where humans translate intent into machine-checkable artifacts, is the connective tissue every layer depends on, yet it lacks four marks of a mature engineering discipline: shared vocabulary, design principles, composability standards, and governance practices. We propose six design principles for Layer 2, from elicitability and composability to adversary-awareness, traceability, and governability, made concrete through worked examples and a reference architecture turning specifications into runtime enforcement, evaluation, and escalation. Existing systems such as Cedar, Constitutional AI, and Open Policy Agent each address a fragment of Layer 2 well and the matrix poorly; treating them as fragments of one shared layer makes composition tractable. As evidence, we introduce CARMA, a Layer 2 prototype for autonomous ETL agents in which one specification drives enforcement, evaluation, and escalation, with every decision traceable to a versioned specification, naming what AI oversight is missing and giving independent teams principles to build the missing pieces so they compose.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑