发表机构
Queensland Technology International School of Engineering, Dalian University of Technology; School of Science and Engineering, The Chinese University of Hong Kong (Shenzhen)(大连理工大学青岛国际工程学院; 香港中文大学(深圳)理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文从强化学习视角研究智能反射面盲波束成形,提出梯度采样方案替代均匀采样,并在真实原型系统中验证其可将平均信噪比提升超过5 dB。
AI 中文摘要
智能反射面(IRS)的波束成形问题长期以来都是从优化角度考虑的,并假设信道状态信息(CSI)是可用的。然而,现实情况是,现有原型很少遵循这种基于模型的方法,因为对于目前的网络协议和硬件而言,信道估计在技术上困难且成本高昂。最近的一个趋势是在没有信道知识的情况下进行盲波束成形。本文从强化学习的角度研究盲波束成形。我们首先表明,现有的盲波束成形方法归结为强化学习背景下贪心算法的一个特例。我们分析了由此产生的累积遗憾,并进一步提出了一个上近似以优化探索概率。此外,我们表明,与现有盲波束成形方法中采用的均匀采样方案相比,梯度采样方案可以提高强化学习的效率。我们通过证明所提出的梯度采样方案是梯度上升的随机近似来进一步验证其收敛性。最后,我们在真实世界的原型系统中物理实现了所提出的方法。我们的现场测试结果表明,与现有的盲波束成形方法相比,所提出的梯度采样在1000个样本下将平均信噪比(SNR)提高了超过5 dB。
英文摘要
The beamforming problem of intelligent reflecting surface (IRS) has been extensively considered from an optimization perspective assuming that channel state information (CSI) is available. However, the reality is that the existing prototypes seldom follow this model-based approach because channel estimation is technically difficult and costly for the network protocols and hardware to date. A recent trend is to perform beamforming blindly without channel knowledge. This work looks at blind beamforming from a reinforcement learning point of view. We first show that the existing blind beamforming method boils down to a special case of the greedy algorithm in the reinforcement learning context. We analyze the resulting cumulative regret, and further propose an upper approximation to facilitate the optimization of the exploration probability. Moreover, we show that a gradient sampling scheme can improve the efficiency of reinforcement learning as compared to the uniform sampling scheme adopted in the existing blind beamforming method. We further verify the convergence of the proposed gradient sampling scheme by showing that it is a stochastic approximation to gradient ascent. Finally, we physically implement the proposed method in a real-world prototype system. Our field test results show that, as compared to the existing blind beamforming method, the proposed gradient sampling boosts the average signal-to-noise ratio (SNR) by more than 5 dB with 1000 samples.