AI 中文总结
本文针对6G开放AI-RAN,首次开展深度强化学习(DRL)综合调查,梳理DRL基础、O-RAN感知框架、DRL应用分类及相关研究方向,填补了DRL方法论与O-RAN部署间的系统关联空白。
AI 中文摘要
向第六代(6G)网络的演进正在将无线接入网(RAN)转变为可编程且智能的控制平台,该平台必须持续适应异构服务、动态环境以及相互竞争的性能目标。开放无线接入网(O-RAN)提供了支持这一转型所需的开放接口、分离式架构和多时间尺度控制环路,而深度强化学习(DRL)为在不确定性下优化序列决策提供了自然框架。然而,现有调查要么广泛探讨O-RAN中的人工智能(AI)和机器学习(ML),要么聚焦于孤立的DRL用例,在DRL方法论、O-RAN架构与运营部署之间的系统关联方面存在空白。据我们所知,本文首次针对面向开放AI-RAN的DRL开展了专门且全面的调查。我们回顾了无模型、基于模型、离线、安全、多智能体、联邦及迁移学习的基础,并提供了一种感知O-RAN的框架,用于通过状态、观测、动作、奖励、约束和时间结构来构建RAN控制问题。我们将DRL应用归类为无线资源管理、移动性管理、干扰控制、流量导向、能效、网络切片、集成感知与通信、安全以及大规模MIMO领域。我们进一步研究了多智能体与联邦协调、基础模型与智能体AI、可信DRL、模拟到真实迁移、持续自适应、资源高效推理以及强化学习运维。最后,我们回顾了实验平台、基准、标准和行业活动,并确定了面向6G开放AI-RAN的样本高效、安全、可扩展、互操作且可部署的DRL控制的研究方向。
英文摘要
The evolution toward sixth-generation (6G) networks is transforming the radio access network (RAN) into a programmable and intelligent control platform that must continuously adapt to heterogeneous services, dynamic environments, and competing performance objectives. Open Radio Access Network (O-RAN) provides the open interfaces, disaggregated architecture, and multi-timescale control loops needed to support this transformation, while deep reinforcement learning (DRL) offers a natural framework for optimizing sequential decisions under uncertainty. However, existing surveys either address artificial intelligence (AI) and machine learning (ML) in O-RAN broadly or focus on isolated DRL use cases, leaving a gap in the systematic connection between DRL methodology, O-RAN architecture, and operational deployment. To the best of our knowledge, this article presents the first dedicated and comprehensive survey of DRL for Open AI-RAN. We review the foundations of model-free, model-based, offline, safe, multi-agent, federated, and transfer learning, and provide an O-RAN-aware framework for formulating RAN control problems through states, observations, actions, rewards, constraints, and temporal structure. We classify DRL applications across radio resource management, mobility management, interference control, traffic steering, energy efficiency, network slicing, integrated sensing and communication, security, and massive MIMO. We further examine multi-agent and federated coordination, foundation models and agentic AI, trustworthy DRL, sim-to-real transfer, continual adaptation, resource-efficient inference, and reinforcement learning operations. Finally, we review experimental platforms, benchmarks, standards, and industry activities, and identify research directions toward sample-efficient, safe, scalable, interoperable, and deployable DRL control for 6G Open AI-RAN.