AI 中文总结
针对异构海上传感器网络单船跟踪问题,提出信息增益引导的强化学习传感器选择框架,通过近端策略优化智能体选传感器,结合贝叶斯跟踪器估计船只状态,经模拟比较,该策略在单传感器激活下性能接近全传感且避免复杂计算。
AI 中文摘要
本文提出了一种信息增益引导的强化学习传感器选择框架,用于异构海上传感器网络中的单船跟踪。该方法受信息论传感器管理启发,通过学习策略在每个决策时刻选择一个与跟踪相关的传感器,而非激活所有传感器或重复进行计算昂贵的在线期望信息增益评估。利用贝叶斯序贯蒙特卡罗跟踪器从噪声测量中估计船只状态,并在非线性和非高斯条件下提供用于调度的置信表示。近端策略优化智能体从塞浦路斯阿伊亚纳帕码头的CMMI智能码头测试床地理参考模拟中部署的五个传感器中选择一个,观察多种特征,奖励由可观测性掩码门控的实现信息增益项定义。最终测试模拟将该框架与随机单传感器选择、同时使用所有传感器的始终开启传感以及之前工作中提出的期望信息增益传感器选择基线进行比较。结果表明,学习到的策略在每个决策时间步仅激活一个传感器的情况下,实现了接近始终开启传感的跟踪性能,同时避免了期望信息增益选择所需的计算昂贵的在线熵搜索。
英文摘要
This paper presents an information-gain-guided reinforcement-learning sensor-selection framework for single-vessel tracking in heterogeneous maritime sensor networks. The proposed approach is motivated by information-theoretic sensor management: instead of activating all sensors or repeatedly performing computationally expensive online expected-information-gain evaluation, a learned policy selects one tracking-relevant sensor at each decision epoch. A Bayesian sequential Monte Carlo tracker estimates the vessel state from noisy measurements and provides a belief representation for scheduling under nonlinear and non-Gaussian conditions. A Proximal Policy Optimization agent selects one of five sensors in a georeferenced simulation of the CMMI Smart Marina testbed at Ayia Napa Marina, Cyprus. The policy is trained on the testbed's actual five-sensor configuration. The agent observes belief-state, detection-history, coverage, sensor-geometry, and realized-information-gain features. The reward is defined as a realized-information-gain term gated by an observability mask. Final-test simulations compare the proposed framework with random single-sensor selection, always-on sensing using all sensors simultaneously, and the expected-information-gain sensor-selection baseline proposed in our previous work. Results show that the learned policy achieves tracking performance close to always-on sensing while activating only one sensor per decision time step and avoiding the computationally expensive online entropy search required by expected-information-gain selection. Additional zero-shot evaluation without retraining on ten moderately perturbed versions of actual layout configuration showed broadly stable tracking, with any increase in positional tracking error remaining below 1 meter across all perturbations.
Comments5 pages, 4 figures, accepted for the IEEE MetroSea 2026 Conference: Special Session 13: Object Detection, Tracking, and Sensor Fusion for Maritime Situational Awareness