发表机构
Software College, Northeastern University; Hohai University(东北大学软件学院; 河海大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多AUV自组织网络目标跟踪的训练不稳定、跟踪性能差问题,提出VGG-MADiffRL算法与MDCA架构,实现了更快收敛、更高跟踪精度与平稳训练动态,具有工程应用价值。
AI 中文摘要
基于多自主水下航行器(AUV)自组织网络的目标跟踪,要求网络化的AUV在受限声学通信、动态拓扑和不确定海洋扰动条件下协同跟踪机动目标。尽管多智能体强化学习(MARL)可通过集中式训练实现去中心化协调,但现有方法存在高维联合状态-动作建模、策略生成对噪声敏感的问题,导致训练不稳定、跟踪性能下降。为解决这些问题,我们提出VGG-MADiffRL(一种值梯度引导多智能体扩散强化学习算法)和MDCA(一种基于扩散的分层控制架构)。利用水下任务特性,我们对声呐探测机制和洋流扰动进行建模,将多AUV自组织网络的协同跟踪问题表述为马尔可夫决策过程(MDP)。所提出的MDCA构成三层闭环控制框架:全局智能控制层、局部在线训练层和物理动作执行层,该结构可实现任务分配、局部决策过程与执行反馈之间的协同优化。在MDCA中,局部在线训练层为策略学习框架;VGG-MADiffRL基于扩散策略构建,在反向去噪过程中融入值梯度以引导动作生成,使生成的动作向更高预期回报方向调整;其采用双值网络进行联合优化与软目标更新,以缓解高估问题和训练振荡,促进更稳定的收敛。实验结果表明,在协同跟踪场景中,VGG-MADiffRL始终实现更快的收敛速度、更高的跟踪精度和更平稳的训练动态,验证了其在动态水下环境中的有效性和实际工程价值。
英文摘要
Multi-AUV ad-hoc network-based target tracking requires networked autonomous underwater vehicles (AUVs) to cooperatively track maneuvering targets under constrained acoustic communication, dynamic topology, and uncertain ocean disturbances. Although multi-agent reinforcement learning (MARL) enables decentralized coordination through centralized training, existing methods suffer from high-dimensional joint state-action modeling, noise-sensitive policy generation, leading to unstable training and degraded tracking. To address these issues, we propose VGG-MADiffRL, a value-gradient-guided multi-agent diffusion RL algorithm, and MDCA, a diffusion?based hierarchical control architecture. Leveraging underwater mission characteristics, we model sonar detection mechanisms and ocean current disturbances, formulating cooperative tracking for multi-AUV ad-hoc networks as an MDP. The proposed MDCA constitutes a three-tier closed-loop control framework: a global intelligent control layer, a local online training layer, and a physical action execution layer. This structure enables synergistic optimization across task allocation, local decision processes, and execution feedback. Within MDCA, the local online training layer is the policy learning framework; VGG-MADiffRL builds on diffusion policies and incorporates value gradients to guide action generation in the reverse denoising process, steering the generated actions towards higher expected returns. It employs twin value networks with joint optimization and soft target updates to mitigate overestimation and training oscillations, promoting more stable convergence. Experimental results show that VGG-MADiffRL consistently achieves faster convergence, higher tracking accuracy, and smoother training dynamics in cooperative tracking scenarios, validating its effectiveness and practical engineering value in dynamic underwater settings.