arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05689cs.AI

迈向AI可信性:为AI控制系统寻找解析可证的向前不变集

Toward AI Trustworthiness: Finding Analytically Proven Forward-Invariant Sets for AI-Controlled Systems

Haoyang Song, Xikun Yang, Qixin Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对AI控制系统的可信性,提出利用可逆神经网络变换状态空间以寻找解析可证的向前不变集,在45个系统中全部成功,优于基线。

中文摘要 AI 辅助

神经网络(NN)控制器越来越多地用于非线性控制系统,但其高度非线性的行为使其难以解释和验证,从而在安全关键和任务关键应用中引发可信性担忧。迈向可认证可信性的关键一步是找到向前不变集(FIS):一个状态空间区域,使得任何从该区域内部开始的轨迹都保持在区域内。如果FIS排除了不安全状态,则对于其中的初始状态可以保证安全性。对于具有固定控制器的给定AI控制系统,寻找解析可证的FIS是困难的。我们提出一个框架,使用可逆神经网络(INN)将原始状态空间变换到潜在空间,在该空间中更可能存在规则形状的FIS。我们训练INN,使得一个优选超矩形候选在潜在空间中变得不变,然后对其进行形式验证。我们证明,每当验证成功时,潜在空间中的候选及其在原始状态空间中的逆变换对应物都是解析可证的FIS。我们在三个代表性控制测试平台上的45个AI控制系统中评估了该方法。我们的方法为所有45个系统找到了经过认证的FIS,而改编的最先进基线则一个也没有找到。该方法在45个系统中的40个上也更快,并且所得FIS的中心大致与领域专家的偏好相符。

英文摘要

Neural-network (NN) controllers are increasingly used in nonlinear control systems, but their highly nonlinear behavior makes them difficult to explain and verify, raising trustworthiness concerns in safety- and mission-critical applications. A key step toward certifiable trustworthiness is to find a Forward-Invariant Set (FIS): a state-space region such that any trajectory starting inside remains inside. If the FIS excludes unsafe states, safety can be guaranteed for initial states within it. Finding an analytically proven FIS for a given AI-controlled system with a fixed controller is difficult. We propose a framework that uses an Invertible Neural Network (INN) to transform the original state space into a latent space where a regular-shaped FIS is more likely to exist. We train the INN so that a preferred hyper-rectangular candidate becomes invariant in the latent space, then formally verify it. We prove that, whenever verification succeeds, both the latent-space candidate and its inverse-transformed counterpart in the original state space are analytically proven FISs. We evaluate the approach on 45 AI-controlled systems across three representative control testbeds. Our method finds certified FISs for all 45 systems, whereas an adapted state-of-the-art baseline finds none. It is also faster on 40 of the 45 systems, and the centers of the resulting FISs roughly match domain-expert preferences.

发表机构

  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑