发表机构
University of Naples Federico II; New York University; University of Southern California; University of Trento; Centre Tecnològic de Telecomunicacions de Catalunya (CTTC)(那不勒斯费德里科二世大学; 纽约大学; 南加州大学; 特伦托大学; 加泰罗尼亚电信技术中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对动态异构网络中延迟敏感信息传输的截止期限约束问题,提出含EC指标与MADRL EC($p^*$)架构的部署导向框架及MGA-RL训练协议,加速演示驱动强化学习以实现可部署网络控制。
AI 中文摘要
在动态异构网络中及时传输延迟敏感信息对下一代(NextG)交互式应用至关重要,但提供严格的端到端(E2E)峰值延迟保证仍是一项开放性挑战。两个障碍限制了学习型网络控制在该场景中的应用:传统基于流量的路由指标虽对通用流量管理极为有效,但并非为捕获流量紧迫性而设计;从头训练的深度强化学习(DRL)控制器存在样本效率低、训练时间长以及早期探索波动大的问题。本文提出一种面向部署的网络控制框架,以解决上述两个障碍。首先,我们提出有效拥塞(EC)这一感知截止期限的指标族,该指标族通过数据包紧迫性量化接口拥塞并主动过滤不可行流量,同时结合统一路径分组(UPG)分布启发式方法以促进稳健的负载均衡;由此产生的策略被嵌入到多智能体深度强化学习有效拥塞($p^*$)(MADRL EC($p^*$))中,这是一种将分布式调度器与基于集中式强化学习的路由器相结合的混合架构。其次,我们提出一种统一训练目标,该目标将现有的策略学习范式——行为克隆、离线强化学习(RL)、在线RL以及离线到在线方案——概括为特例,它结合了实时奖励项、预收集奖励项和策略模仿项。从该目标出发,我们推导了模型引导退火强化学习(MGA-RL)协议,该协议以深度确定性策略梯度(DDPG)为骨干实例化,是一种面向部署的、由演示驱动的训练方法,它概括了传统的离线到在线(O2O)方案,其中来自轻量级[...]
英文摘要
Timely delivery of delay-sensitive information over dynamic, heterogeneous networks is essential for NextG interactive applications, yet providing strict End-to-End (E2E) peak latency guarantees remains an open challenge. Two obstacles limit the adoption of learning-based network control in this setting: traditional volume-based routing metrics, while highly effective for general traffic management, are not designed to capture traffic urgency; and Deep Reinforcement Learning (DRL) controllers trained from scratch suffer from sample inefficiency, long training times, and early-stage exploration volatility. This paper introduces a deployment-focused network control framework that addresses both obstacles. First, we present Effective Congestion (EC), a deadline-aware metric family that quantifies interface congestion by packet urgency and proactively filters non-viable traffic, coupled with a Uniform Path Grouping (UPG) distribution heuristic promoting robust load-balancing; the resulting policies are embedded into Multi-Agent Deep Reinforcement Learning Effective Congestion ($p^*$) (MADRL EC ($p^*$)), a hybrid architecture combining a distributed scheduler with a centralized RL-based router. Second, we introduce a unified training objective that generalizes existing policy-learning paradigms---behavioral cloning, offline Reinforcement Learning (RL), online RL, and offline-to-online schemes---as special cases, combining a live-reward term, a pre-collected-reward term, and a policy-imitation term. From this objective, we derive the Model-Guided Annealed Reinforcement Learning (MGA-RL) protocol, instantiated on a Deep Deterministic Policy Gradient (DDPG) backbone: a deployment-oriented, demonstration-driven training approach that generalizes conventional Offline-to-Online (O2O) schemes, in which trajectories from a lightweight [...]