arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于梯度聚类BS采样的快速收敛元强化学习用于边缘缓存

Fast-Convergent Meta-RL via Gradient-Clustered BS Sampling for Edge Caching

Farnaz Niknia, Ping Wang

arXiv 2609.16370首次发表:更新:

AI 中文总结

针对边缘缓存中元RL训练采样方差高的问题,提出基于梯度聚类的基站采样方法,降低元梯度方差并加速收敛。

AI 中文摘要

无线边缘缓存网络通常由许多独立的基站(BS)组成,每个基站面临各自的请求速率和内容流行度分布。在每个基站从零开始训练一个强化学习(RL)缓存智能体,迫使每个智能体通过缓慢的试错过程重新学习一个在网络中结构相同的决策问题。元强化学习通过学习一个共享的初始化来消除这种冗余,该初始化能在少量本地更新后适应任何基站;然而,元训练本身成为大规模场景下的瓶颈:元梯度必须在每次元迭代中从一小部分基站子集估计,而均匀随机采样该子集会得到高方差的估计,现有元RL缓存框架未解决此问题。本文提出一种针对独立、不重叠基站的缓存元强化学习框架,直接针对该瓶颈。每个基站运行一个本地近端策略优化(PPO)智能体,将其表述为关于内容流行度、大小、生命周期和重要性的半马尔可夫决策过程(SMDP),同时通过模型无关元学习(MAML)风格的循环学习共享元策略。为扩展元训练规模并加速收敛,我们引入基于梯度的聚类,根据本地梯度相似性对基站分组,并在每次元迭代中按比例从每个簇中采样。通过梯度方差的方差分析(ANOVA)式分解,我们证明在基站异质性下,该策略产生的元梯度估计器严格低于均匀随机采样的方差。

英文摘要

Wireless edge caching networks typically consist of many independent Base Stations (BSs), each facing its own request rate and content popularity profile. Training a Reinforcement Learning (RL) caching agent from scratch at every BS forces each agent to relearn, through slow trial and error, a decision problem that is structurally identical across the network. Meta-reinforcement learning removes this redundancy by learning a shared initialization that adapts to any BS in a few local updates; however, meta-training itself becomes the bottleneck at scale: the meta-gradient must be estimated from a small subset of BSs at each meta-iteration, and sampling this subset uniformly at random yields a high-variance estimate, an issue existing meta-RL caching frameworks leave unaddressed. This paper proposes a meta-reinforcement learning framework for caching across independent, non-overlapping BSs that directly targets this bottleneck. Each BS runs a local Proximal Policy Optimization (PPO) agent, formulated as a Semi-Markov Decision Process (SMDP) over content popularity, size, lifetime, and importance, while a shared meta-policy is learned via a Model-Agnostic Meta-Learning (MAML)-style loop. To scale meta-training and accelerate convergence, we introduce gradient-based clustering, which groups BSs by local gradient similarity and draws from every cluster, in proportion to its size, at each meta-iteration. We prove, via an Analysis of Variance (ANOVA)-style decomposition of gradient variance, that this strategy yields a strictly lower-variance meta-gradient estimator than uniform random sampling under BS heterogeneity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑