arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16950math.OC

混合信道上远程状态估计的最优调度

Optimal Scheduling for Remote State Estimation over Hybrid Channels

Manali Dutta, Rahul Singh, Shalabh Bhatnagar

首次发表
浏览论文内容

中文总结 AI 辅助

研究混合信道上远程状态估计的最优调度,将其建模为MDP,刻画了最优策略的阈值结构,在系统参数未知时提出AC学习算法,数值结果显示该算法能有效学习策略结构,性能接近最优策略。

中文摘要 AI 辅助

我们研究了在具有两种异构通信信道(快速但不可靠的信道和慢速但可靠的信道)的网络上进行远程状态估计的最优调度问题。为捕捉分组丢失中的时间相关性,将不可靠信道建模为吉尔伯特 - 埃利奥特(GE)信道。远程估计设置包括一个源、一个传感器和一个远程估计器。源按离散时间自回归(AR)过程演化,传感器每次决定使用快速不可靠信道还是慢速可靠信道。我们将传感器面临的调度问题表述为具有连续状态空间的马尔可夫决策过程(MDP),并考虑最小化无限时域平均成本准则,成本包括估计误差平方和传输能耗。我们建立了最优平稳策略的存在性,刻画了最优策略的结构,表明其具有关于估计误差的阈值结构。当系统参数未知时,提出一种利用最优策略阈值结构的演员 - 评论家(AC)学习算法。数值结果表明,所提出的AC算法能有效学习策略结构,性能接近使用相对值迭代(RVI)计算的最优策略。

英文摘要

We study optimal scheduling for remote state estimation over a network with two heterogeneous communication channels: a fast but unreliable channel and a slow but reliable channel. To capture temporal correlations in packet losses, we model the unreliable channel as a Gilbert-Elliott (GE) channel. The remote estimation setup consists of a source, a sensor, and a remote estimator. The source evolves as a discrete-time autoregressive (AR) process, and the sensor decides at each time whether to use the fast unreliable channel or the slow reliable channel. We formulate the scheduling problem faced by the sensor as a Markov decision process (MDP) with a continuous state-space and consider minimizing the infinite horizon average cost criterion, where the cost consists of the squared estimation error and the transmission energy consumed. We establish the existence of an optimal stationary policy. We then characterize the structure of an optimal policy, and show that it has a threshold structure with respect to the estimation error. An optimal policy chooses from amongst the two channels based on whether the error exceeds certain thresholds, where the threshold value depends upon the GE channel state. When the system parameters are unknown, we propose an actor-critic (AC) learning algorithm that exploits the threshold structure of an optimal policy. Numerical results demonstrate that the proposed AC algorithm learns the policy structure effectively and achieves performance close to that of the optimal policy computed using the relative value iteration (RVI).

↑