AI 中文总结
针对未知统计特性的时变马尔可夫信道调度问题,提出Online-MGF策略,可最大化有限时间内总吞吐量,在开/关信道下推导Whittle索引闭式,性能收敛快。
AI 中文摘要
我们考虑下行无线通信网络中基站(BS)向多个用户发送数据的无线调度问题,其中信道统计特性未知。调度性能高度依赖于准确的信道状态信息(CSI),而获取这类信息往往成本高昂。在本文中,CSI仅在每次调度传输后通过确认(ACK)/非确认(NACK)反馈获取。由于无线信道资源有限,无法同时调度所有用户进行传输,因此最新观测到的CSI可能已过时。传统利用过时CSI解决调度问题的方法是使用置信状态,该状态通过信道的时间相关统计特性计算,但信道统计特性通常未知,导致置信状态可能不可数,该方法变得不可行。本文为无线调度问题引入了一种新的充分统计量,为此我们用信道状态信息年龄(AoCSI)表征CSI的陈旧性,并证明最新观测到的CSI及其AoCSI是用于做出调度决策的历史的充分统计量,从而能够减少在线学习的状态空间。我们的目标是开发一种在线调度算法,在满足信道资源约束的有限时间范围内最大化所有用户的期望总吞吐量,所构建的问题是一个 restless 多臂老虎机(RMAB)问题。我们开发了在线最大增益优先(Online-MGF)策略,该策略在回合数上实现次线性遗憾。对于开/关信道的特殊情况,我们能够证明其索引性并推导Whittle索引的闭式表达式。数值结果表明,Online-MGF策略在极少的回合内就收敛到具有已知统计特性的MGF策略和Whittle索引策略。
英文摘要
We consider a wireless scheduling problem in downlink wireless networks with unknown channel statistics, where a Base Station (BS) sends data to multiple users. The scheduling performance relies heavily on accurate Channel State Information (CSI), which is often costly to acquire. In this paper, CSI is obtained from ACK/NACK feedback, only after each scheduled transmission. Due to limited wireless channel resources, all users cannot be scheduled for transmission simultaneously. Hence, the most recently observed CSI can be outdated. The traditional approach to solve scheduling problems using outdated CSI is to utilize belief states, which are calculated using the time correlation statistics of channels. However, channel statistics are often unknown; consequently, belief states can be uncountable and this approach becomes infeasible. In this paper, we introduce a new sufficient statistics for the wireless scheduling problem. Towards this effort, we characterize the CSI staleness by the Age of Channel State Information (AoCSI) and show that the latest observed CSI and its AoCSI is a sufficient statistic of the history to make the scheduling decisions. Accordingly, we are able to reduce the state space for online learning. Our goal is to develop an online scheduling algorithm that maximizes the expected sum throughput of all users over a finite time-horizon while satisfying a channel resource constraint. The formulated problem is a Restless Multi-armed Bandit (RMAB). We develop an online Maximum Gain First (Online-MGF) policy, which achieves sub-linear regret on the number of episodes. For a special case of ON/OFF channels, we are able to prove indexability and derive a closed-form expression of the Whittle index. Numerical results demonstrate that the Online-MGF policy converges to MGF and Whittle index policies with known statistics within a very few episodes.
Comments19 pages, 8 figures. This manuscript is currently under review in IEEE/ACM Transactions on Networking