用于EXL-50U实验校准模拟环境中X点靶磁构型控制的优势级聚合强化学习
Advantage-level Aggregation Reinforcement Learning for X-point Target Magnetic Configuration Control in an EXL-50U Experiment-Calibrated Simulation Environment
另 7 家 · 查看机构详情
- Beijing ENN Fusion Energy Science and Technology Co., Ltd.(北京新奥融合能源科技有限公司)
- Beijing Key Laboratory of High Magnetic Field Spherical Torus Fusion Energy(北京强磁场球形托卡马克聚变能源重点实验室)
- Hebei Key Laboratory of Compact Fusion(河北省紧凑型聚变重点实验室)
- Dalian University of Technology(大连理工大学)
- School of Physics, Dalian University of Technology(大连理工大学物理学院)
- Key Laboratory of Materials Modification by Beams of the Ministry of Education(教育部材料表面与界面工程重点实验室)
- School of Nuclear Science and Engineering, East China University of Technology(东华理工大学核科学与工程学院)
- East China University of Technology(东华理工大学)
- School of Mechanics and Engineering Science, Peking University(北京大学力学与工程科学学院)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究针对EXL-50U实验校准模拟环境中的X点靶磁构型控制问题,提出AdvA算法优化PPO策略,实现了X点通量控制精度的大幅提升,为后续实时验证奠定基础。
中文摘要 AI 辅助
管理偏滤器热负荷是紧凑型高功率托卡马克的核心挑战。为增加局部通量膨胀并将耗散体积与芯部解耦,EHL-2采用了X点靶(XPT)偏滤器,这要求次级X点保持在偏滤器支腿上,其位移会破坏拓扑结构和排气几何形状。当前包括EXL-50U放电在内的实验依赖于预计算的前馈波形与全局量的PID控制回路,由于缺乏针对次级零位的专用闭环反馈,XPT操作可重复但非常规。我们将XPT反馈建模为在针对EXL-50U第13906次放电校准的自由边界环境中的多目标强化学习(RL)控制问题。为解决等离子体电流、形状和零位约束之间的强耦合——其中奖励标量化会导致特定目标的时间信用崩溃——我们开发了优势聚合(AdvA)方法。AdvA在感知最差目标的非线性标量化之前保留目标层面的时间信用,并为策略更新引入残差校正。我们在标称操作、测量不确定性和未见过的初始平衡条件下,将AdvA-PPO与Reward-PPO及前馈加PID基线进行评估。在500毫秒的滚动测试中,AdvA-PPO将平均最差通道得分从0.23提升至0.81,较Reward-PPO降低约20倍的X点通量RMSE;在组合测量不确定性下,它是唯一能完成时间范围同时保持可用XPT形状的学习型控制器。多初始化微调使单个AdvA-PPO策略能在偏滤器和限制器初始平衡条件下完成全时间范围操作,这些结果为未来在EXL-50U上进行实时XPT验证提供了基于模拟的基础。
英文摘要
Managing divertor heat loads is a central challenge for compact, high-power tokamaks. To increase local flux expansion and decouple the dissipation volume from the core, EHL-2 adopts the X-point target (XPT) divertor. This requires the secondary X-point to remain on the divertor leg; displacement degrades the topology and exhaust geometry. Current experiments, including EXL-50U discharges, rely on precomputed feedforward waveforms with PID loops on global quantities. Lacking dedicated closed-loop feedback for the secondary null, XPT operation is repeatable but not routine. We formulate XPT feedback as a multi-objective reinforcement learning (RL) control problem in a free-boundary environment calibrated to EXL-50U discharge #13906. To address strong coupling among plasma current, shape, and null constraints - where reward scalarisation collapses objective-specific temporal credit - we develop Advantage Aggregation (AdvA). AdvA preserves objective-wise temporal credit before worst-objective-aware nonlinear scalarisation and introduces a residual correction to policy updates. AdvA-PPO is evaluated against Reward-PPO and a feedforward-plus-PID baseline under nominal operation, measurement uncertainties, and unseen initial equilibria. On a 500 ms rollout, AdvA-PPO raises the mean worst-channel score from 0.23 to 0.81 over Reward-PPO, reducing X-point flux RMSE by ~20x. Under combined measurement uncertainties, it is the only learned controller completing the horizon while retaining a usable XPT shape. Multi-initialization fine-tuning enables a single AdvA-PPO policy to complete full-horizon operation across divertor and limiter initial equilibria. These results provide a simulation-based foundation for future real-time XPT validation on EXL-50U.