发表机构
Department of Mechanical and Nuclear Engineering, Tennessee Technological University(机械与核工程系,田纳西技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对未知物理噪声参数的可靠性最优POMDP,提出证书感知有源消歧框架,定义噪声信息价值(VoNI),通过界定VoNI表明其在不同区域的决策价值,双层决策者依后验触发探测,模拟显示该方法能更早检测噪声跳跃等,降低误差和减少探测动作。
AI 中文摘要
有限可靠性表示(FRR)可证明在具有已知物理噪声底的部分观测系统中,单元恒定策略何时足以进行可靠决策。然而在实际中,传感和执行噪声可能是潜在且依赖上下文的。本文针对未知物理噪声参数theta = (sigma_y, sigma_u) 开发了一个证书感知有源消歧框架,固定sigma_u可得到仅传感器情况。我们将噪声信息价值(VoNI)定义为使用根据当前估计而非实际噪声参数校准的可靠性覆盖所导致的预期超额FRR证书差距。通过动作值模型不匹配和FRR半径膨胀来界定VoNI,表明在FRR证书对theta不敏感的子交叉区域中,噪声估计的决策价值较低,但当后验不确定性会使当前覆盖无效时变得有价值。一个双层决策者使用从创新统计、执行残差或其他在线估计器获得的theta后验,并仅在不确定性威胁到FRR证书时触发诊断探测。我们还将VoNI解释为用于潜在传感 -执行区域消歧的高级有限POMDP的可处理、证书感知近似。在平稳、可识别和持续激励的区域中,我们建立了后验一致性以及诱导策略损失收敛到FRR近似下限。基于EKF的创新残差的闭环UGV模拟显示,在50次蒙特卡罗试验中,比后验熵探索更早检测到突然的传感噪声跳跃,具有更低的漂移跟踪误差和显著更少的探测动作。
英文摘要
Finite Reliability Representations (FRR) certify when a cell-constant policy is sufficient for reliable decision-making in a partially observed system with a known physical noise floor. In practice, however, sensing and execution noise can be latent and context-dependent. This paper develops a certificate-aware active disambiguation framework for an unknown physical noise parameter theta = (sigma_y, sigma_u), with the sensor-only case obtained by fixing sigma_u. We define the Value of Noise Information (VoNI) as the expected excess FRR certificate gap caused by using a reliability cover calibrated to the current estimate rather than to the realized noise parameter. We bound VoNI using action-value model mismatch and FRR radius inflation, showing that noise estimation has low decision value in sub-crossover regimes where the FRR certificate is insensitive to theta, but becomes valuable when posterior uncertainty can invalidate the current cover. A bi-level decision maker uses a posterior over theta, obtained from innovation statistics, execution residuals, or another online estimator, and triggers diagnostic probing only when uncertainty threatens the FRR certificate. We also interpret VoNI as a tractable, certificate-aware approximation to a high-level finite POMDP for latent sensing-execution regime disambiguation. Under stationary, identifiable, and persistently exciting regimes, we establish posterior consistency and convergence of the induced policy loss to the FRR approximation floor. Closed-loop UGV simulations with EKF-based innovation residuals show earlier detection of abrupt sensing-noise jumps, lower drift-tracking error, and substantially fewer probing actions than posterior-entropy exploration over 50 Monte Carlo trials.
Comments13 pages, 4 figures, 1 table. Simulation code is available in the accompanying public GitHub repository