发表机构
Linköping University(林雪平大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对重叠传感器覆盖下的多源远程估计,提出约束马尔可夫决策过程模型,证明资源函数秩至多K,并给出基于拉格朗日对偶的求解方法,数值验证了策略随机化与预算交互。
AI 中文摘要
我们研究了由K个具有重叠覆盖的传感器观测的多个有限状态马尔可夫源的语义感知远程估计问题。传感器共享时分多址上行链路,并在传输可靠性、传输延迟和传输预算方面有所不同。在每个时隙中,调度器联合选择一个源及其一个监测传感器,或保持空闲,以在全局和每个传感器的传输频率约束下最小化长期平均执行误差成本。我们将此问题建模为有限平均成本约束马尔可夫决策过程。我们证明传输资源函数的秩至多为K,尽管全局约束仍可能限制可行域。因此,拉格朗日函数仅依赖于K个有效传输成本,并且最优约束解可以用至多K+1个确定性策略循环类分量表示。我们进一步刻画了分段仿射凹拉格朗日值函数,并在显式有界的乘子集上推导出投影对偶次梯度上升法。数值结果展示了值函数结构、代表性实例中策略随机化的必要性,以及全局与每个传感器传输预算之间的相互作用。
英文摘要
We study semantic-aware remote estimation of multiple finite-state Markov sources observed by K sensors with overlapping coverage. The sensors share a time-division multiple-access uplink and differ in transmission reliability, delivery delay, and transmission budget. In each slot, the scheduler jointly selects a source and one of its monitoring sensors, or remains idle, to minimize the long-run average cost of actuation error subject to global and per-sensor transmission-frequency constraints. We formulate this problem as a finite average-cost constrained Markov decision process. We show that the transmission resource functions have rank at most K, although the global constraint may still restrict the feasible region. Consequently, the Lagrangian depends only on K effective transmission costs, and an optimal constrained solution can be represented using at most K+1 deterministic policy-recurrent-class components. We further characterize the piecewise-affine concave Lagrangian value function and derive projected dual subgradient ascent over an explicitly bounded multiplier set. Numerical results illustrate the value-function structure, the need for policy randomization in a representative instance, and the interaction between global and per-sensor transmission budgets.