发表机构
Shanghai University; Edith Cowan University(上海大学; 埃迪斯科文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对无线边缘网络中LLM推理服务的有效吞吐量最大化问题,提出两阶段可屏蔽近端策略优化算法,仿真显示其系统奖励提升33.3%至87.5%且有效吞吐量最高。
AI 中文摘要
本文提出一种新颖的两阶段可屏蔽近端策略优化(TP-MPPO)算法,该算法针对无线边缘网络中大型语言模型(LLM)推理服务,在严格遵守服务水平目标(SLO)的前提下,最大化计入请求吞吐量的系统有效吞吐量。在TP-MPPO的第一阶段,通过带有动作屏蔽机制的MPPO优化任务卸载决策,有效避免探索无效动作并缩小动作空间;在第二阶段,推导上行链路带宽分配的闭式解,设计下行链路带宽分配的贪心算法,为下一轮MPPO提供即时奖励。两个阶段交替执行直至收敛。仿真结果表明,与基准方法相比,TP-MPPO可将系统奖励提升33.3%至87.5%,并实现最高的有效吞吐量。
英文摘要
This paper presents a novel two-phase maskable proximal policy optimization (TP-MPPO) algorithm, which maximizes the system goodput counting request throughput with strict service level objective (SLO) compliance for large language model (LLM) inference services in wireless edge networks. In the first phase of TP-MPPO, we optimize the task offloading decisions by MPPO with action masking mechanism, effectively avoiding exploring invalid actions and reducing the action space. In the second phase, closed-form solutions are derived for uplink bandwidth allocation; a greedy algorithm is designed for downlink bandwidth allocation to provide immediate rewards for the MPPO in the next round. The two stages alternate till convergence. Simulation results demonstrate that TP-MPPO can improve the system reward by 33.3%--87.5% compared to its benchmarks and achieve the highest goodput.