arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

A-MADiff:面向移动AIGC网络中感知内存的任务编排的注意力引导多智能体深度强化学习与扩散策略

A-MADiff: Attention-Guided Multi-Agent DRL with Diffusion Policies for Memory-Aware Task Orchestration in Mobile AIGC Networks

Chongzhi Wu, Zhengtao Li, Jiawen Kang, Jinbo Wen, Xiaohuan Li, Maomao Zhang, Ekram Hossain

arXiv 2608.29255首次发表:更新:

发表机构

School of Automation, Guangdong University of Technology; City University of Hong Kong; School of Information and Communication, Guilin University of Electronic Technology; Anhui Engineering Research Center for Agricultural Product Quality Safety Digital Intelligence, Fuyang Normal University; School of Physics and Electronic Engineering, Fuyang Normal University; University of Manitoba(广东工业大学自动化学院; 香港城市大学; 桂林电子科技大学信息与通信学院; 阜阳师范大学安徽省农产品质量安全数字智能工程研究中心; 阜阳师范大学物理与电子工程学院; 曼尼托巴大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对移动AIGC网络中AIGC推理任务导致GPU内存耗尽的问题,本文提出A-MADiff算法,通过协作式多智能体框架建模Dec-POMDP,采用扩散策略与注意力引导机制,显著提升了累积奖励。

AI 中文摘要

人工智能生成内容(AIGC)服务采用生成式AI(GenAI)模型自动生成多样化内容。移动AIGC网络将GenAI模型部署在位于边缘的AIGC服务提供商(ASP)上,为移动用户提供低延迟、个性化的AIGC服务。然而,AIGC推理任务通常会占用GPU内存直至任务完成,这会导致提供服务的ASP出现GPU内存耗尽的情况,进而触发内存不足故障,而不仅仅是增加服务延迟。现有关于AIGC任务编排的研究在很大程度上忽略了GPU内存可行性约束。为解决该问题,本文开发了一种协作式多智能体编排框架,其中每个边缘节点配备一个调度智能体,用于将任务路由到本地ASP或相邻边缘节点。由于调度智能体仅基于本地观测做出决策,而对等卸载会耦合它们的资源状态和长期效用,因此本文将编排过程建模为协作式分散部分可观测马尔可夫决策过程(Dec-POMDP)。为求解该Dec-POMDP,本文提出了一种注意力引导多智能体深度强化学习算法,该算法采用扩散策略(A-MADiff),遵循集中式训练与分散式执行范式。A-MADiff采用基于扩散的分散式行动者,生成可行编排动作的多模态偏好,并采用注意力引导的集中式评论者,在GPU内存异构性下从跨智能体状态估计每个智能体的价值。数值结果表明,与最先进的基线相比,A-MADiff显著提高了累积奖励。

英文摘要

Artificial Intelligence-Generated Content (AIGC) services employ Generative AI (GenAI) models to automatically generate diverse content. Mobile AIGC networks host GenAI models on edge-located AIGC Service Providers (ASPs) to deliver low-latency and personalized AIGC services for mobile users. However, AIGC inference tasks typically occupy GPU memory until task completion, causing GPU memory exhaustion at serving ASPs and triggering out-of-memory failures rather than merely increasing service latency. Existing studies on AIGC task orchestration have largely overlooked GPU memory feasibility constraints. To address this issue, we develop a cooperative multi-agent orchestration framework, in which each edge node is equipped with a scheduling agent to route tasks to local ASPs or neighboring edge nodes. Since scheduling agents make decisions based only on local observations, while peer offloading couples their resource states and long-term utilities, we formulate the orchestration process as a cooperative Decentralized Partially Observable Markov Decision Process (Dec-POMDP). To solve the Dec-POMDP, we propose an \underline{A}ttention-guided \underline{M}ulti-\underline{A}gent deep reinforcement learning algorithm with \underline{Diff}usion policies (A-MADiff) under the centralized training with a decentralized execution paradigm. A-MADiff employs diffusion-based decentralized actors to generate multi-modal preferences over feasible orchestration actions, and an attention-guided centralized critic to estimate per-agent values from cross-agent states under GPU memory heterogeneity. Numerical results demonstrate that A-MADiff significantly improves the cumulative reward over the state-of-the-art baseline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑