AI 中文总结
针对未来无线系统面临的挑战,提出Fused-CPRO方法,将分配策略融合可学习目标策略、源策略和领域知识策略,通过约束随机逐次凸逼近处理非凸问题,经仿真验证该方法能提升经验性能且收敛更快。
AI 中文摘要
深度强化学习因其无模型适应性而被广泛用于无线资源分配。然而,在线探索成本高昂,随机初始化策略可能在收集到足够数据之前违反长期约束。未来无线系统面临动态流量、波动信道条件和严格能效要求,需要能快速学习且环境交互最少的算法。我们开发了Fused-CPRO,一种知识融合的约束策略重用优化方法。它将分配策略构建为可学习目标策略、相关场景源策略和专家规则领域知识策略的混合,在约束马尔可夫决策过程中联合优化目标策略和重用概率。异构先验融合加速收敛并增强鲁棒性。约束随机逐次凸逼近处理非凸目标和约束,基于混合离线-在线数据训练的评论家通过重用预收集经验提高样本效率。我们证明其几乎必然收敛到Karush-Kuhn-Tucker点。在延迟约束多用户多输入多输出功率控制和克拉美-罗下界约束多输入多输出集成感知与通信波束成形的仿真表明,Fused-CPRO提高了经验性能且收敛速度比代表性基线快得多。
英文摘要
Deep reinforcement learning (DRL) has been widely adopted for wireless resource allocation due to its model-free adaptability. However, online exploration is costly, as randomly initialized policies may violate long-term constraints before sufficient data are collected. Future wireless systems must cope with increasingly dynamic traffic, fluctuating channel conditions, and stringent energy efficiency requirements, demanding algorithms that can learn quickly with minimal environment interactions to reduce both energy consumption and signaling overhead. We develop Fused-CPRO, a knowledge-fused constrained policy reuse optimization method addressing these challenges. Fused-CPRO constructs the allocation policy as a mixture of a learnable target policy, source policies from related scenarios, and domain-knowledge (DK) policies from expert rules, jointly optimizing the target policy and reuse probabilities under a constrained Markov decision process (CMDP). This fusion of heterogeneous priors accelerates convergence and enhances robustness. Constrained stochastic successive convex approximation (CSSCA) handles non-convex objectives and constraints, while a critic trained from mixed offline-online data improves sample efficiency by reusing pre-collected experience. We prove almost-sure convergence to a Karush-Kuhn-Tucker (KKT) point. Simulations on delay-constrained multi-user multiple-input multiple-output (MU-MIMO) power control and Cramer-Rao bound (CRB)-constrained multiple-input multiple-output integrated sensing and communication (MIMO-ISAC) beamforming demonstrate that Fused-CPRO improves empirical performance and converges substantially faster than representative baselines.
Comments19 pages, 5 figures, 1 algorithm. Includes appendices with detailed convergence proofs