arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于学习的、感知软截止期限的Transformer增强PPO用于LLM推理的协作式移动边缘计算(MEC)

Learning-Based Collaborative MEC for LLM Inference with Soft-Deadline Awareness via Transformer-Enhanced PPO

Ngoc Hung Nguyen, Bjorn Landfeldt

arXiv 2608.02031首次发表:更新:

发表机构

Lund University(隆德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对软截止期限下LLM推理的协作MEC问题,提出Transformer增强PPO框架,通过捕捉时间与跨服务器交互优化任务迁移,在任务完成率和系统效率上优于传统方法。

AI 中文摘要

本文研究软截止期限约束下用于大语言模型(LLM)推理的协作式移动边缘计算(MEC)服务器。在该系统中,为提升服务质量,计算需在截止期限内完成,但因任务或子任务间存在依赖关系,任何截止期限的错过都可能对整个请求造成灾难性后果。为此,本研究提出一种具有受限灵活性的扩展截止期限机制。主要挑战在于,在严格延迟约束下处理大规模计算,同时限制允许的截止期限扩展次数,尤其在每个请求内存在任务依赖关系的情况下。为应对这些挑战,我们开发了Transformer增强的近端策略优化(PPO)框架,该框架可实现MEC服务器间的高效协作。所提方法旨在最大化在截止期限内完成的任务数量,同时最小化截止期限扩展的使用次数。通过捕捉时间依赖关系和跨服务器交互,Transformer改进了任务迁移的决策。仿真结果表明,所提方法在任务完成率和整体系统效率方面显著优于传统PPO及基于启发式的方法。

英文摘要

This paper investigates collaborative mobile edge computing (MEC) servers for large language model (LLM) inference under soft deadline constraints. In this system, to improve the quality of service, computations are expected to be completed within their deadlines. However, due to dependencies among tasks or subtasks, any missed deadline can lead to catastrophic consequences for the entire request. In this context, this work proposes an extended deadline mechanism with constrained flexibility. The main challenges lie in handling large-scale computations under strict latency constraints while limiting the number of allowable deadline extensions, especially in the presence of task dependencies within each request. To tackle these challenges, we develop a transformer-enhanced proximal policy optimization (PPO) framework that enables efficient collaboration among MEC servers. The proposed approach aims to maximize the number of tasks completed within their deadlines while minimizing the use of deadline extensions. By capturing temporal dependencies and cross-server interactions, the transformer improves decision-making for task migration. Simulation results demonstrate that the proposed method significantly outperforms conventional PPO and heuristic-based approaches in terms of task completion rate and overall system efficiency.

Comments7 pages, 5 pages

Journal ref2026 IEEE GLOBECOM SELECTED AREAS IN COMMUNICATIONS: CLOUD/EDGE COMPUTING AND NETWORKING

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑