发表机构
University of Washington, Tacoma(华盛顿大学塔科马分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出BRaVeS框架,通过不变锚点、深度感知访问及李雅普诺夫有界共识机制,在认知风险增加时降低自主性,并在模拟中验证了有限步收敛与无安全违规,为智能体AI自动化提供可验证的安全保障。
AI 中文摘要
由智能体AI赋能的自动化在高风险环境中不能仅依靠概率推理来安全部署。一个反复出现的风险是认知漂移:随着推理的深入,系统行为可能偏离领域专家对安全运行的约束。本文提出了BRaVeS,一个被称为“可辩护的下一代推理系统”(DNRS)的有界推理与安全治理框架。BRaVeS将领域专家定义的约束编码为不变锚点,提出MoDA风格(混合深度注意力)的深度感知访问作为在推理过程中保持这些锚点可见的候选机制,并使用状态层级(SMARtAutonomy)在认知风险增加时降低自主性。为了形式化有界恢复,我们引入了李雅普诺夫有界共识框架(LBCF),该框架将连续的认知风险信号映射到有限的K袋抽象中,并应用屏蔽状态转换,强制实现李雅普诺夫式能量下降或将系统引导至人工介入的终端状态。形式化收敛结果适用于在固定阈值和可行屏蔽假设下的有限LBCF抽象;它并不证明完整连续神经激活空间的安全性。我们通过使用HAI 22.04工业控制系统时间序列数据(带有合成噪声和传感器退化机制)的离散事件蒙特卡洛模拟来评估该框架。在测试的参数分组策略和阈值下,LBCF过程实现了有限步收敛且无安全防护违规。这些结果提供了基于模拟的证据,表明在所陈述的抽象下可以强制实施有界治理行为,同时推动了未来在部署的Transformer实现、实时人在回路验证以及更广泛的对抗性设置方面的工作。
英文摘要
Agentic AI-enabled automation cannot be safely deployed in high-stakes environments on probabilistic reasoning alone. A recurring risk is epistemic drift: as reasoning deepens, system behavior may move away from subject-matter-expert constraints for safe operation. This paper presents BRaVeS, a bounded reasoning and safety-governance framework termed the Defensible Next-Gen Reasoning System (DNRS). BRaVeS encodes SME-defined constraints as invariant anchors, proposes MoDA-Style (Mixture of Depths Attention) depth-aware access as a candidate mechanism for keeping these anchors visible during inference, and uses a state hierarchy (SMARtAutonomy) to reduce autonomy as epistemic risk increases. To formalize bounded recovery, we introduce the Lyapunov-Bounded Consensus Framework (LBCF), which maps continuous epistemic-risk signals into a finite K-bag abstraction and applies shielded state transitions that enforce Lyapunov-style energy descent or route the system to a human-mediated terminal state. The formal convergence result applies to the finite LBCF abstraction under fixed thresholds and feasible-shield assumptions; it does not prove safety of the full continuous neural activation space. We evaluate the framework through a discrete event Monte Carlo simulation using HAI 22.04 industrial-control-system time-series data with synthetic noise and sensor-degradation regimes. Across the tested parameter-grouping strategies and thresholds, the LBCF process achieved finite-step convergence and no safety-guard violations. These results provide simulation-based evidence that bounded governance behavior can be enforced under the stated abstraction, while motivating future work on deployed transformer implementations, live human-in-the-loop validation, and broader adversarial settings.
CommentsThis paper has been accepted and will appear in the Journal of Intelligent and Robotic Systems (10.1007/s10846-026-02467-w)