AI 中文总结
研究资源受限、动态无线网络环境中基于意图的无人机轨迹设计问题,提出基于逆强化学习的数字孪生驱动解决方案,相比传统方法性能损失大幅减少,相比标准无人机网络性能提升2.5倍。
AI 中文摘要
本文研究了在资源受限、动态无线网络环境中运行的基于意图的无人机轨迹设计问题。在该模型中,无人机作为辅助基站在地面用户集群间导航,提供按需上行数据接入。数字孪生(DT)系统在中央服务器上创建物理无线网络环境的虚拟表示以模拟和预测相关变化,在此情况下DBS轨迹也应调整。然而,DT系统在无法保证获取DBS意图(即服务优先级)的情况下建议调整轨迹,因为该意图随时间演变且因DBS与DT服务器间的间歇性连接无法及时更新到DT系统。这种调整被视为一个优化问题,目标是找到使DBS服务的优先用户比例最大化的轨迹。为解决未知DBS意图和不可预测的环境变化问题,提出了一种基于逆强化学习(IRL)的DT驱动解决方案。仿真结果表明,与基于传统强化学习的机载DBS控制相比,该解决方案能提供近实时、近最优的轨迹调整,在环境变化时性能损失减少约85%。与标准无人机网络相比,DT框架还将无人机网络性能提高了2.5倍,在标准网络中DBS运行时存在错误和延迟的环境感知。
英文摘要
In this paper, the problem of the trajectory design for an intent-based drone operating in resource-constrained, dynamic wireless network environments is studied. In the considered model, the drone acts as a supplementary base station that navigates among ground user clusters to provide on-demand uplink data access. Given its intended application (e.g traffic monitoring), the drone base station (DBS) prioritizes serving certain clusters (e.g. high-risk highway sections). A digital twin (DT) system, hosted on a central server, creates a virtual representation of the physical wireless network environment to simulate and predict related changes, in which case the DBS trajectory should also be adjusted. Then, the DT system suggests adjustments to DBS trajectories without guaranteed access to the underlying DBS intent (i.e., service priorities), as this intent evolves over time and cannot be updated to the DT system in a timely manner due to intermittent connectivity between the DBS and the DT server. Such adjustment is posed as an optimization problem whose goal is to find the trajectories with which the fraction of prioritized users served by the DBS is maximized. To solve this problem under unknown DBS intent and unpredictable environment changes, an inverse reinforcement learning (IRL) based DT actuation solution is proposed. Simulation results demonstrate that the proposed solution provides near-real-time, near-optimal trajectory adjustment, with approximately 85\% less performance loss across environmental changes, compared to traditional reinforcement learning based on-board DBS control. The DT framework also enhances drone network performance by up to 2.5 times, compared to standard drone networks where a DBS operates with its erroneous and delayed environmental sensing.
Comments12 pages, 11 figures