发表机构
University of Calgary; Ericsson Canada Inc.(卡尔加里大学; 加拿大爱立信公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RRM-GPT自回归框架,将无线电资源管理决策视为序列生成,通过预训练与后训练实现跨部署和跨功能复用,并在5G NR案例中验证单模型生成完整调度授权。
AI 中文摘要
基于学习的无线电资源管理(RRM)模型通常针对单一功能和部署场景而构建,因此每种新设置都需要重复开发流程。然而,RRM决策具有共同的结构:每个决策由标准定义且相互依赖的字段组合而成,其值根据网络状态进行选择。我们提出RRM-GPT,一种用于RRM基础模型的自回归框架,该框架以语言模型生成文本的方式生成这些决策。编码器将异构网络观测映射为统一的令牌表示,解码器逐字段输出决策,每个字段均以网络状态和已提交字段为条件。在未标注的网络日志上进行预训练,使模型学习什么构成有效决策以及控制器如何在有效决策中进行选择;随后通过模仿学习或强化学习进行后训练,使其适应特定部署的运营商目标。该框架针对两种复用形式:跨部署复用的功能特定模型,以及跨RRM功能共享的模型,该模型将其相互依赖的决策生成为一个序列。在5G新空口(NR)案例研究中,我们证明单个模型能够生成涵盖用户选择、定时、链路自适应、资源分配和控制信令的完整调度授权。该模型捕获授权字段间的依赖关系,并将学习到的行为迁移至未见场景而无需适应。
英文摘要
Learning-based models for radio resource management (RRM) are typically built for a single function and deployment, so each new setting repeats the development pipeline. RRM decisions, however, share a common structure: each is assembled from interdependent fields, defined by the standard, whose values are selected in view of the network state. We propose RRM-GPT, an autoregressive framework for RRM foundation models that generate these decisions as a language model generates text. An encoder maps heterogeneous network observations into a common token representation, and a decoder emits the decision one field at a time, each conditioned on the network state and the fields already committed. Pretraining on unannotated network logs teaches the model what makes a decision valid and how controllers choose among valid decisions; post-training then adapts it to deployment-specific operator objectives through imitation or reinforcement learning. The framework targets two forms of reuse: a function-specific model reused across deployments, and a model shared across RRM functions that generates their interdependent decisions as one sequence. In a 5G New Radio (NR) case study, we demonstrate that a single model generates complete scheduling grants spanning user selection, timing, link adaptation, resource allocation, and control signaling. The model captures dependencies among grant fields and transfers learned behavior to an unseen scenario without adaptation.