arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

主动式非对角可重构智能表面赋能异构边缘计算:一种分布式强化学习方法

Active Beyond-Diagonal RIS Empowered Heterogeneous Edge Computing: A Distributional Reinforcement Learning Approach

Tianyu Pang, Hongyu Li

arXiv 2607.13160首次发表:更新:

发表机构

Internet of Things Thrust, The Hong Kong University of Science and Technology (Guangzhou)(物联网 thrust,香港科学与技术大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究互易主动BD-RIS辅助异构MEC的能量感知卸载与资源分配问题,基于分布式软演员-评论家算法改进版DSAC-T构建端到端联合优化框架,相比其他算法,实现了更好的能量-延迟奖励、可行性比率及更快在线决策时间

AI 中文摘要

主动式非对角可重构智能表面(BD-RIS)能实现混合发射与反射模式,为异构移动边缘计算(MEC)系统中的阻塞感知上行卸载提供方案。但实际混合模式主动BD-RIS由互易设备实现,会产生跨扇区能量泄漏。本文研究互易主动BD-RIS辅助异构MEC的能量感知卸载与资源分配,提出基于分布式软演员-评论家算法改进版DSAC-T的端到端联合优化框架,相比其他算法,DSAC-T实现了最佳能量-延迟奖励、最高可行性比率81.67%及快速在线决策时间。

英文摘要

Active beyond-diagonal reconfigurable intelligent surfaces (BD-RISs) enables hybrid transmitting and reflecting mode to achieve effective signal amplification and full-space coverage, thus providing a promising solution for blockage-aware uplink offloading in heterogeneous mobile edge computing (MEC) systems. However, practical hybrid mode active BD-RIS are realized by reciprocal devices, which inherently generate cross-sector energy leakage that will reshape the system-level energy-latency tradeoff. This paper studies energy-aware offloading and resource allocation for reciprocal active BD-RIS-assisted heterogeneous MEC, where offloading decisions, CPU/GPU computation allocation, transmit powers, receive processing, and active BD-RIS are tightly coupled. The resulting problem is a high-dimensional mixed integer nonconvex problem and is difficult to solve efficiently by conventional per-instance optimization. To address this challenge, we develop an end-to-end joint optimization framework based on a refined version of the distributional soft actor--critic algorithm, named as DSAC-T. By modeling return distributions rather than only expected values, DSAC-T improves policy stability under reward heterogeneity and feasibility-boundary sensitivity. Compared with other baseline algorithms, DSAC-T achieves the best energy-latency reward, the highest feasibility ratio of 81.67%, and a fast online decision time of 0.0267 s per scenario.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑