AI 中文总结
针对边缘云网络负载不平衡致延迟性能下降问题,提出联合服务放置等的JSCP问题。利用决策动态差异分解问题,构建双时间尺度多层深度强化学习框架2T-MDRL-LA,引入潜行动作表示,有效优化相关内容,相比传统方法性能更优且收敛更快。
AI 中文摘要
在分层边缘云计算(HECC)系统中,动态任务到达和异构资源下边缘层与云层间的负载不平衡会降低延迟性能,导致严重排队延迟和资源利用效率低下。为应对这一挑战,研究联合服务放置、计算委托和功率控制(JSCP)问题以最小化平均端到端延迟。该问题因离散与连续变量强耦合而成混合整数非凸且NP难优化问题。为实现易处理的优化和稳定系统适应,利用决策动态固有差异将问题分解为长期系统配置和短期资源分配子问题。基于此提出具有潜行动作空间的双时间尺度多层深度强化学习框架(2T-MDRL-LA)来联合优化服务放置、用户关联、计算委托、任务卸载和用户发射功率。引入基于变分自编码器的潜行动作表示以有效压缩高维组合动作空间。仿真结果表明该框架能有效适应动态网络条件,与分支定界解相比性能接近最优,平均端到端延迟最多降低20.8%,资源利用率提高13%,且比传统近端策略优化收敛快约50%。
英文摘要
Load imbalance across edge and cloud layers degrades latency performance in hierarchical edge-cloud computing (HECC) systems under dynamic task arrivals and heterogeneous resources, leading to severe queuing delays and inefficient resource utilization. To address this challenge, we study a joint service placement, computational delegation, and power control (JSCP) problem to minimize the average end-to-end (e2e) latency. The resulting JSCP problem is a mixed-integer nonconvex and NP-hard optimization problem due to the strong coupling between discrete and continuous variables. To enable tractable optimization and stable system adaptation, we exploit the inherent difference in decision dynamics and decompose the problem into long-term system configuration and short-term resource allocation subproblems. Based on this formulation, we propose a two-timescale multi-layer deep reinforcement learning framework with a latent action space (2T-MDRL-LA) to jointly optimize service placement, user association, computational delegation, task offloading, and user transmit power. A latent action representation based on a variational autoencoder is introduced to efficiently compress the high-dimensional combinatorial action space. Simulation results demonstrate that the proposed framework effectively adapts to dynamic network conditions and achieves near-optimal performance compared to branch-and-bound solutions. It achieves up to a 20.8% reduction in average e2e latency and a 13% improvement in resource utilization over the scheme without the computational delegation, while converging approximately 50% faster than conventional proximal policy optimization.
CommentsSumitted for publication (14 pages)