arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26068cs.NIcs.LG

多层网络可行性几何上的可微策略传输

Differentiable Policy Transport over Multi-Layer Network Feasibility Geometry

Zuyuan Zhang, Zeyu Fang, Mahdi Imani, Nathaniel D. Bastian, Tian Lan

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出 NFG-RL,通过传输理论将异构网络约束建模为可行性几何,实现可微策略传输,在无线边缘环境中显著提升可行效用、降低违规和时延。

中文摘要 AI 辅助

基于学习的控制正日益成为网络操作自动化的核心。然而,学习到的策略必须满足跨层约束,涉及干扰、功率-速率耦合、流守恒、服务链、容量、时延和可靠性。现有方法通常仅考虑该几何结构的子集,且仅间接处理,例如通过奖励惩罚、拉格朗日乘子或事后修复。本文提出网络可行性几何强化学习(NFG-RL),通过传输理论和残差包含关系 $\bphi_{\mathfrak{N}}(x,a)\in\cK_{\mathfrak{N}}$ 对耦合约束建模,将执行策略定义为原型策略通过可行性传输映射的推前(pushforward)。NFG-RL 将异构约束编译为类型化残差块,并通过可微变分算子传输原型动作,使活动约束塑造执行、探索和演员梯度。我们的分析表明,精确传输可实现几乎必然可行的执行,而活动约束将探索收缩到可行切空间上。进一步,分析确立了批评者倾斜传输相对于普通投影具有非负一阶增益,并将背压调度恢复为提升漂移残差的梯度。在两个基于公共轨迹条件的无线边缘替代环境中,NFG-RL 在每个环境中相比最强的非 NFG 方法将可行效用提高了 37.5--41.5%,将原始动作违规减少了 48.5--60.8%,并将 P99 时延降低了 57.0--75.5%,优于一系列优化和学习基线。

英文摘要

Learning-based control is increasingly central to automating network operations. A learned policy, however, must satisfy cross-layer constraints on interference, power-rate coupling, flow conservation, service chains, capacity, latency, and reliability. Existing methods typically account for only a subset of this geometry and only indirectly, e.g., through reward penalties, Lagrange multipliers, or post-hoc repairs. This paper proposes \emph{Network Feasibility Geometry Reinforcement Learning} (NFG-RL), which models coupled constraints via transport theory and the residual inclusion $\bphi_{\mathfrak{N}}(x,a)\in\cK_{\mathfrak{N}}$, defining the executed policy as the pushforward of a proto-policy through a feasibility-transport map. NFG-RL compiles heterogeneous constraints into typed residual blocks and transports proto-actions through a differentiable variational operator, letting active constraints shape execution, exploration, and actor gradients. Our analysis shows that exact transport yields almost-sure feasible execution, while active constraints contract exploration onto the feasible tangent space. It further establishes a nonnegative first-order gain from critic-tilted transport over plain projection and recovers backpressure scheduling as the gradient of a lifted drift residual. In two public-trace-conditioned wireless-edge surrogate environments, NFG-RL improves feasible utility by \textbf{37.5--41.5\%} over the strongest non-NFG method in each environment, reduces raw-action violation by \textbf{48.5--60.8\%}, and lowers P99 delay by \textbf{57.0--75.5\%}, outperforming a range of optimization and learning baselines.

发表机构

  • The George Washington University(乔治华盛顿大学)
  • Northeastern University(东北大学)
  • Johns Hopkins University(约翰斯·霍普金斯大学)
  • Syracuse University(雪城大学)

机构由 AI 辅助整理,请以论文原文为准。

↑