arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通用马尔可夫源的拉式状态估计的错误信息年龄

Age of Incorrect Information for Pull-Based State Estimation of General Markov Sources

Marco Zanni, Mohamad Assaad, Touraj Soleymani

arXiv 2608.13248首次发表:更新:

AI 中文总结

该研究针对通用马尔可夫源的拉式状态估计,以错误信息年龄为目标,建立马尔可夫决策过程,提出多源调度的近似Whittle索引策略,其启发式策略性能接近最优且计算量小。

AI 中文摘要

我们研究了任意多状态马尔可夫源的拉式远程状态估计,同时考虑了信息的新鲜度和正确性属性。为此,我们针对错误信息年龄(AoII)制定了一个折扣优化问题,并将其表示为最大后验概率(MAP)估计下的联合源-AoII信念马尔可夫决策过程(MDP)。随后,我们利用模型的信息结构,证明每个可达信念由最后一次成功观测到的源状态以及自该观测以来经过的时隙数表示。为了进行数值计算,我们将未成功的持续时间截断到有限水平,并推导了明确的误差界和选择截断参数的准则。对于可靠链路,我们证明最优策略可由等待时间的查找表表示。对于不可靠链路,我们提出了一种持久策略并推导了可计算的性能界。我们还证明MAP估计在有限时隙后会稳定。为进一步降低内存需求,我们引入了具有早期平稳切换的混合估计器,并推导了由此产生的性能差异的可计算界。最后,我们将框架扩展到多个源,将调度问题表述为 restless 多臂老虎机,建立了可索引性的充分条件,并开发了基于插值的近似Whittle索引策略。我们的数值结果说明了单源最优策略的结构,评估了多源策略的性能,并验证了所提出的启发式策略在大幅减少计算工作量的同时,非常接近最优解。

英文摘要

We study pull-based remote state estimation of an arbitrary, multi-state Markov source while accounting for both freshness and correctness attributes of information. To that end, we formulate a discounted optimization problem in terms of the age of incorrect information (AoII), and express it as a joint source-AoII belief Markov decision process (MDP) under maximum a posteriori (MAP) estimation. We then exploit the information structure of the model and prove that every reachable belief is represented by the last successfully observed source state and the number of time slots elapsed since that observation. For numerical computation, we truncate the elapsed no-success duration at a finite level and derive an explicit error bound and a criterion for selecting the truncation parameter. For reliable links, we show that an optimal policy can be represented by a look-up table of waiting times. For unreliable links, we propose a persistent policy and derive computable performance bounds. We also show that the MAP estimate stabilizes after a finite number of time slots. To further reduce memory requirements, we introduce a hybrid estimator with an early stationary switch and derive a computable bound on the resulting difference in performance. Finally, we extend the framework to multiple sources, formulate the scheduling problem as a restless multi-armed bandit, establish a sufficient condition for indexability, and develop an approximate Whittle index policy based on interpolation. Our numerical results illustrate the structure of the optimal single-source policy, evaluate the performance of the multi-source policies, and verify that the proposed heuristic policies closely approach the optimal solution while substantially reducing computational efforts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑