发表机构
School of Electrical Engineering and Computer Science, University of Ottawa(渥太华大学电气工程与计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对6G无线网络的多任务移动网络优化挑战,提出WiSDoM框架,结合DT与MoE架构,在多种网络配置下训练,性能优于对比方法,可高效适应未见场景。
AI 中文摘要
新兴的6G无线网络预计将在多样化的部署场景中运行,网络拓扑、用户移动性、流量需求和无线条件的变化对传统无线资源管理(RRM)的可扩展性构成了挑战。尽管离线强化学习(RL)方法已展现出强大的决策能力,但由于优化目标相互冲突且模型专业化程度有限,学习一个在异构无线环境中表现一致的单一策略仍然困难。这些挑战在协作多点(CoMP)传输中尤为突出,选择最优服务小区组合需要在不断变化的网络条件下进行序列决策。本文提出了带混合专家的无线稀疏决策Transformer(WiSDoM),这是一个用于自适应多小区选择的稀疏多任务离线RL框架。WiSDoM将决策Transformer(DTs)与混合专家(MoE)架构相结合,该架构可根据任务特征动态激活专门的专家。这种MoE机制在不按比例增加推理成本的情况下提高了模型容量,减轻了负迁移,并实现了专家在不同任务间的专业化。WiSDoM在涵盖多种基站和用户设备密度、移动性水平以及调度器策略的多样化网络配置上进行联合训练。实验结果表明,WiSDoM始终优于启发式方法、单任务模型和传统多任务DTs,将体验质量(QoE)提升了高达55%,且在推理过程中仅激活其密集对应模型约三分之一的参数。此外,WiSDoM表现出强大的任务泛化能力,无需重新训练或微调即可通过少样本提示高效适应未见的无线场景。
英文摘要
Emerging 6G wireless networks are expected to operate across diverse deployment scenarios, where variations in network topology, user mobility, traffic demand, and radio conditions challenge the scalability of conventional radio resource management (RRM). While offline reinforcement learning (RL) methods have demonstrated strong decision-making capabilities, learning a single policy that performs consistently across heterogeneous wireless environments remains difficult due to conflicting optimization objectives and limited model specialization. These challenges become particularly pronounced in coordinated multipoint (CoMP) transmission, where selecting the optimal serving-cell combination requires sequential decision-making under evolving network conditions. This paper presents the Wireless Sparse Decision Transformer with Mixture of Experts (WiSDoM), a sparse multi-task offline RL framework for adaptive multi-cell selection. WiSDoM combines Decision Transformers (DTs) with a Mixture-of-Experts (MoE) architecture that dynamically activates specialized experts according to task characteristics. This MoE mechanism improves model capacity without proportionally increasing inference cost, mitigates negative transfer, and enables expert specialization across tasks. WiSDoM is trained jointly on diverse network configurations spanning multiple base station and user equipment densities, mobility levels, and scheduler policies. Experimental results show that WiSDoM consistently outperforms heuristic methods, single-task models, and conventional multi-task DTs, improving quality of experience (QoE) by up to 55% while activating approximately one-third of the parameters of its dense counterpart during inference. Furthermore, WiSDoM exhibits strong task generalization and efficiently adapts to unseen wireless scenarios through few-shot prompting without retraining or fine-tuning.
Comments13 pages, 11 figures, submitted to IEEE for possible publication