AI 中文总结
研究自由空间光通信系统中光学组件精确位置控制问题,采用深度确定性策略梯度(DDPG)智能体调谐PID控制器,经实验表明其在动态定位任务中有潜力,不同目标轨迹下效果有差异,需多模式优化。
AI 中文摘要
在自由空间光通信系统中,光学组件的精确位置对于保持光束对准至关重要。本文研究了用于控制在光学系统焦平面中移动光纤末端的光学偏转器的级联位置和速度PID控制器的强化学习辅助调谐。一个深度确定性策略梯度(DDPG)智能体通过与物理实验台交互来调整六个PID系数。该实验台支持高达12kHz的目标坐标更新,智能体与受控设备相距约200km并通过UDP交换数据。经过5000次训练后,选择两组固定系数并与手动调谐的基线进行比较。对于伪随机目标轨迹,最佳的强化学习调谐集将径向定位误差范围从119减小到82,降幅为31%,其标准差从15减小到12。对于恒定零目标,强化学习调谐集未改善径向误差范围。结果证明了DDPG在动态定位任务中进行实验性PID调谐的潜力,并表明需要多模式优化以在不同操作条件下实现一致性能。
英文摘要
Accurate positioning of optical components is essential for maintaining beam alignment in free-space optical (FSO) communication systems. This work investigates reinforcement-learning-assisted tuning of cascaded position and velocity PID controllers for an optical deflector that moves the end of an optical fiber in the focal plane of an optical system. A Deep Deterministic Policy Gradient (DDPG) agent adjusts six PID coefficients through interaction with a physical experimental stand. The stand supports target-coordinate updates of up to $12$ kHz, while the agent and the controlled device are located approximately $200$ km apart and exchange data over UDP. After $5000$ training sessions, two fixed coefficient sets are selected and compared with a manually tuned baseline. For a pseudo-random target trajectory, the best RL-tuned set reduces the range of the radial positioning error from $119$ to $82$, corresponding to a $31\%$ reduction, and decreases its standard deviation from $15$ to $12$. For a constant zero target, the RL-tuned sets do not improve the radial error range. The results demonstrate the potential of DDPG for experimental PID tuning in dynamic positioning tasks and indicate the need for multi-regime optimization to achieve consistent performance under different operating conditions.