arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

风险感知的在线共形状态探测

Risk-Aware Online Conformal State Probing

Pietro Talli, Petar Popovski, Osvaldo Simeone

arXiv 2609.25889首次发表:更新:

发表机构

Institute for Intelligent Networked Systems, Northeastern University London; Aalborg University(东北大学伦敦智能网络系统研究所; 奥尔堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出在线共形状态探测(OCSP)策略,在无分布假设下认证最坏情况可靠性,控制遗漏查询误差并最小化探测率,适用于预训练控制策略且无需重训。

AI 中文摘要

基于人工智能的自主智能体通常部署在数据中心,必须从机器人或边缘设备获取状态信息,以便做出明智的控制决策。在安全关键场景中,管理状态的不确定性尤为重要,因为平均情况下的保证是不够的。在此背景下,我们研究了一个序贯决策过程,该过程在给定任意状态预测模型的情况下,联合决定采取哪些动作以及何时进行探测。我们提出了在线共形状态探测(OCSP),一种动作和探测策略,它在不依赖分布假设的情况下认证最坏情况下的可靠性水平。OCSP被设计为可证明地控制遗漏查询误差(MQE),即探测本会有益的实例比例,同时最小化探测率。OCSP可以应用于现有的预训练基于值的控制策略,无需重新训练或微调。我们通过数值模拟验证了OCSP,以验证理论保证并评估性能权衡作为状态预测器校准的函数。

英文摘要

AI-based autonomous agents, typically hosted at data centers, must acquire state information from robots or edge devices in order to issue informed control decisions. Managing uncertainty about the state is particularly consequential in safety-critical settings, in which average-case guarantees are insufficient. In this context, we study a sequential decision maker process that jointly decides which actions to take and when to probe given access to an arbitrary state prediction model. We propose online conformal state probing (OCSP), an action and probing policy that certifies worst-case reliability levels without relying on distributional assumptions. OCSP is designed to provably control the missed query error (MQE), i.e., the fraction of instances where probing would have been beneficial, while minimizing the probing rate. OCSP can be applied to existing pre-trained value-based control policies without requiring retraining or fine-tuning. We validate OCSP through numerical simulations to verify theoretical guarantees and to assess performance trade-offs as a function of the calibration of the state predictor.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑