AI 中文总结
该研究针对异构应急服务的无人机辅助应急通信场景,提出MA-HEAD-Net框架,通过自适应规则引导的多智能体深度强化学习优化联合调度,实现了更低的信息年龄(AoI)与更高的策略形成效率。
AI 中文摘要
在灾后场景中,无人机(UAV)对于建立应急通信网络至关重要。对于时间敏感的救援任务而言,信息的新鲜度至关重要,因为基于过时数据做出的决策可能会导致无效的控制行动。本文研究了异构应急服务场景下,无人机辅助应急通信中的信息年龄(Age of Information, AoI)最小化问题。我们使用马尔可夫调制泊松过程对突发的数据包到达进行建模,并采用有限块长理论来捕捉传输时长、数据包完成情况与AoI演化之间的耦合关系。为了平衡延迟容忍型长数据包传输与紧急短数据包响应,我们提出了一种嵌入微时隙的调度机制,该机制包含自适应检查点间隔选择功能。我们将无人机轨迹控制、用户调度以及检查点间隔选择的联合优化问题建模为多智能体决策问题,并开发了MA-HEAD-Net,这是一种自适应规则引导的多智能体深度强化学习框架。MA-HEAD-Net将通信领域的规则先验融入到门控多头策略中,其中自适应门控器会针对不同子任务调节规则先验与学习策略逻辑值的贡献。该策略与门控组件在多智能体近端策略优化的框架下进行联合优化。仿真结果表明,与代表性的多智能体深度强化学习基线方法相比,MA-HEAD-Net提升了策略形成效率,并且在动态无人机辅助应急通信场景中,相较于基于学习的方法和启发式方法,它实现了更低的AoI。
英文摘要
In post-disaster scenarios, unmanned aerial vehicles (UAVs) are critical for establishing emergency communication networks. For time-critical rescue missions, information freshness is crucial because decisions based on outdated data may lead to ineffective control actions. This paper investigates age of information (AoI) minimization for UAV-assisted emergency communications with heterogeneous emergency services. We model bursty packet arrivals using a Markov-modulated Poisson process and adopt finite blocklength theory to capture the coupling among transmission duration, packet completion, and AoI evolution. To balance delay-tolerant long-packet transmission and urgent short-packet response, we propose a mini-slot-embedded scheduling mechanism with adaptive checkpoint-interval selection. We formulate the joint optimization of UAV trajectory control, user scheduling, and checkpoint-interval selection as a multi-agent decision problem, and develop MA-HEAD-Net, an adaptive rule-guided multi-agent deep reinforcement learning framework. MA-HEAD-Net incorporates communication-domain rule priors into a gated multi-head policy, where adaptive gates regulate the contributions of rule-prior and learned-policy logits for different subtasks. The policy and gating components are jointly optimized under multi-agent proximal policy optimization. Simulation results show that MA-HEAD-Net improves policy-formation efficiency compared with representative multi-agent deep reinforcement learning baselines and achieves lower AoI than both learning-based and heuristic methods in dynamic UAV-assisted emergency communication scenarios.
Journal refIEEE Transactions on Cognitive Communications and Networking, vol. 12, pp. 9962-9978, 2026