多用户MIMO系统中基于RL引导的自动编码器切换的系统感知自适应CSI反馈
System-Aware Adaptive CSI Feedback via RL-Guided Autoencoder Switching in Multi-User MIMO System
浏览论文内容
中文总结 AI 辅助
针对mMIMO系统中信道状态信息反馈问题,提出基于强化学习驱动控制框架,在预训练多速率自动编码器组上运行,依信道条件和性能指标选压缩率,引入系统感知奖励公式,相比传统方案提升频谱和反馈效率,降低反馈成本等。
中文摘要 AI 辅助
本文提出了一种用于大规模多输入多输出(mMIMO)系统的系统感知自适应信道状态信息(CSI)反馈框架,旨在动态优化重建保真度和信令开销之间的权衡。基于深度学习的自动编码器(AE)虽能显著压缩CSI,但传统固定比率方案无法有效适应非平稳信道条件。为此,我们开发了一种强化学习(RL)驱动的控制框架,该框架在一组预训练的多速率AE上运行,每个AE对应不同的压缩率(CR)。在每个时间步,集中式RL智能体根据观察到的信道条件和系统性能指标为每个用户选择最合适的CR。与传统以均方误差(MSE)为中心的设计不同,我们引入了一种系统感知奖励公式,该公式通过信号与干扰加噪声比(SINR)、反馈开销约束以及模型适应的计算成本共同考虑频谱效率。在高维延迟域CSI数据集上的仿真结果表明,所提出的RL引导框架有效地平衡了开销-精度权衡,并适应动态信道环境。与固定压缩方案和自适应基线相比,该方法提高了频谱效率和反馈效率,同时保持了适度的计算和内存占用。在不同数量的用户和所有考虑的基线上进行平均,所提出的RL框架将CSI反馈成本降低了53.4%以上,将平均下行链路总和速率提高了53.64%,并将归一化均方误差(NMSE)降低了22.38%。这些结果证明了其在动态无线条件下实现更高效的速率-精度-反馈权衡的能力。
英文摘要
This paper proposes a system-aware adaptive channel state information (CSI) feedback framework for massive multiple-input multiple-output (mMIMO) systems, aiming to dynamically optimize the trade-off between reconstruction fidelity and signaling overhead. While deep learning-based autoencoders (AEs) have enabled significant CSI compression, conventional fixed-ratio schemes fail to adapt effectively to non-stationary channel conditions. To address this limitation, we develop a reinforcement learning (RL)-driven control framework that operates over a bank of pretrained multi-rate AEs, each corresponding to a distinct compression ratio (CR). At each time step, a centralized RL agent selects the most suitable CR for each user based on observed channel conditions and system performance indicators. Distinct from conventional mean squared error (MSE)-centric designs, we introduce a system-aware reward formulation that jointly accounts for spectral efficiency via signal-to-interference-plus-noise ratio (SINR), feedback overhead constraints, and the computational cost of model adaptation. Simulation results on high-dimensional delay-domain CSI datasets demonstrate that the proposed RL-guided framework effectively balances the overhead-accuracy tradeoff and adapts to dynamic channel environments. The proposed method improves spectral efficiency and feedback efficiency compared with fixed compression schemes and adaptive baselines, while maintaining a modest computational and memory footprint. Averaged over different numbers of users and across all considered baselines, the proposed RL framework reduces the CSI feedback cost by more than 53.4%, improves the average downlink sum rate by 53.64%, and reduces the NMSE by 22.38%. These results demonstrate its ability to achieve a more efficient rate-accuracy-feedback tradeoff under dynamic wireless conditions.