发表机构
National Institute of Technology Rourkela; Polytechnique Montréal(印度国家技术研究所鲁尔克拉分校; 蒙特利尔理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对OTFS-NOMA框架中STAR-RIS重构的难题,提出SAC-BSE深度强化学习算法优化其相位与能量分裂系数,可快速收敛并提升和速率性能。
AI 中文摘要
本文研究由同时发射与反射可重构智能表面(STAR-RIS)辅助的正交时频空间(OTFS)与非正交多址接入(NOMA)技术组成的下行通信框架。此类框架中的延迟多普勒移动性使得传统交替优化方法无法适用于每个相干间隔的重构。为缓解这些问题,将STAR-RIS相移与能量分裂设计建模为带约束的非凸和速率最大化问题,采用闭式最大比传输波束成形与固定NOMA功率分配。为规避每个间隔的重新优化负担,采用深度强化学习(DRL)方法,通过单次前向传播将观测到的信道实现映射到STAR-RIS配置。具体而言,提出Beta空间软演员评论家(SAC-BSE),一种最大熵DRL智能体。仿真中,每个STAR-RIS分支有2个NOMA复用用户,结果证实其收敛迅速,在128倍用户速度范围内将和速率下降限制在约10%,且随发射功率与STAR-RIS单元数增加,相比仅OTFS、仅NOMA、仅STAR-RIS、固定分裂及模式切换基线均产生稳定增益。
英文摘要
This paper considers a downlink communication framework comprising a simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS)-aided by orthogonal time frequency space (OTFS) and non-orthogonal multiple access (NOMA) technologies. Further, delay-Doppler mobility in such frameworks renders classical alternating optimization impractical for per-coherence interval reconfiguration. To mitigate such issues, the STAR-RIS phase-shift and energy-splitting design is formulated as a constrained, non-convex sum-rate maximization problem with closed-form maximum ratio transmission beamforming and fixed NOMA power allocation. To circumvent the per-interval re-optimization burden, a deep reinforcement learning (DRL) approach is adopted that maps observed channel realizations to STAR-RIS configurations through a single forward pass. Specifically, Beta-Space Soft Actor-Critic (SAC-BSE), a maximum entropy DRL agent, is proposed. Simulation results, with two NOMA-multiplexed users on each STAR-RIS branch, confirm rapid convergence, limit the sum-rate degradation to roughly 10\% across a 128-fold user-speed range, and yield consistent gains over OTFS-only, NOMA-only, STAR-RIS-only, fixed-split, and mode-switching baselines as transmit power and the number of STAR-RIS elements increase.
CommentsAccepted in IEEE ANTS 2026