基于多智能体强化学习的可重构智能表面序列拍卖:物理层安全分析
MARL-Based Sequential RIS Auctions: A Physical-Layer Security Analysis
浏览论文内容
中文总结 AI 辅助
本文针对合法接收者与窃听者竞争RIS资源的问题,提出基于MARL的序列RIS拍卖框架,采用MADDPG方法求解马尔可夫博弈,使合法接收者单位成本保密速率接近理想上限,优于随机和固定策略。
中文摘要 AI 辅助
可重构智能表面(RIS)通过智能配置其反射元件,在增强覆盖范围、频谱效率和通信安全方面具有巨大潜力。当RIS由中立运营商所有时,这些元件可作为合法接收者和窃听者竞争的资源。本文研究此类竞争并评估其对合法接收者物理层安全性能的影响。为建模该竞争,我们开发了序列RIS拍卖(SRA)框架,其中每轮通过第一价格密封拍卖机制拍卖一组RIS元件,每个竞标者根据可实现速率增益和剩余预算提交竞标。随后,我们通过指定状态、动作、奖励和状态转移,将序列竞标过程建模为马尔可夫博弈。为求解该博弈,我们提出了基于多智能体强化学习(MARL)的多竞标者深度确定性策略梯度(MADDPG)方法,采用集中式训练与分布式执行(CTDE)模式,使合法接收者和窃听者能够学习竞标策略以最大化其长期经济盈余。数值结果表明,在考虑的窃听者竞标策略下,基于强化学习的策略使合法接收者实现最高的单位成本保密速率,优于随机策略和固定策略,且接近理想物理层上限。
英文摘要
Reconfigurable intelligent surfaces (RISs) hold great potential to enhance coverage, spectral efficiency, and communication security by intelligently configuring their reflecting elements. When owned by a neutral RIS operator, these elements can be offered as resources for which legitimate receivers and eavesdroppers compete. This paper investigates such competition and evaluates its impact on the physical-layer security performance of legitimate receivers. To model the competition, we develop a sequential RIS auction (SRA) framework, in which a bundle of RIS elements is auctioned in each round through a first-price sealed-bid mechanism, with each bidder submitting its bid based on the achievable rate gain and remaining budget. We then formulate the sequential bidding process as a Markov game by specifying its states, actions, rewards, and state transitions. To solve the game, we propose a multi-bidder deep deterministic policy gradient (MADDPG)-based multi-bidder reinforcement learning (MARL) approach under centralized training and decentralized execution (CTDE), enabling legitimate receivers and eavesdroppers to learn bidding strategies that maximize their long-term economic surplus. Numerical results show that, under the considered eavesdropper bidding strategies, the RL-based strategy enables legitimate receivers to achieve the highest secrecy rate per unit cost, outperforming random and fixed strategies and approaching the ideal physical-layer upper bound.