arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

改进型长谷川-若菜系统中湍流转变的强化学习控制

Reinforcement-learning control of turbulence transition in the modified Hasegawa-Wakatani system

Luning Sun, Ben Zhu, Xin-Yang Liu, Deepak Akhare, Jian-Xun Wang

arXiv 2608.27845首次发表:更新:

AI 中文总结

该研究将强化学习方法应用于改进型长谷川-若菜系统的湍流-带状流双向控制,通过耦合CNN智能体与JAX求解器,在湍流抑制和逆带状流破坏任务中均取得优于基线的效果,证明了强化学习用于等离子体湍流控制的可行性。

AI 中文摘要

等离子体湍流的控制仍是磁约束聚变研究中一项长期存在的挑战。本文探索一种强化学习(RL)方法,用于对改进型长谷川-若菜系统(静电漂移波湍流的最小模型)中的湍流-带状流转变进行双向控制。在该模型中,弱带状流阻尼会抑制原本寿命很长的带状结构,赋予其有限寿命,并恢复有限控制区间内可重复转变所需的驱动-阻尼平衡;同时,致动通过空间分布的高斯源场施加,且致动成本受时间加权预算约束。随后,等离子体模型通过原生GPU的JAX求解器与基于CNN的软 Actor-Critic 及双延迟确定性策略梯度智能体耦合,该求解器针对快速在线训练进行了高度优化。在湍流抑制任务中,学习到的感知预算策略在给定消耗预算下实现了最低的时间积分湍流通量,在所有测试的未见初始条件下均优于恒定和线性递减基线。在逆带状流破坏任务中,智能体发现了一种上下反对称的致动模式,该模式会诱导径向E×B对流,破坏带状结构并维持湍流状态。物理启发的暖缓冲初始化促进了这一发现,因为仅随机探索难以在庞大的动作空间中找到狭窄的最优流形。这些结果表明,强化学习作为非线性等离子体动力学(如湍流控制)的实用轨迹优化器是可行的,并凸显了此类应用中物理指导的重要性。

英文摘要

Control of plasma turbulence remains a long-standing challenge in magnetically confined fusion research. Here, we explore a reinforcement-learning (RL) approach for bidirectional control of the turbulence-zonal-flow transition in the modified Hasegawa-Wakatani system, a minimal model of electrostatic drift-wave turbulence. Within our model, a weak zonal drag damps the otherwise long-lived zonal structures, giving them a finite lifetime and restoring the drive-damping balance required for repeatable transitions within finite control episodes; meanwhile, actuation is applied through a spatially distributed Gaussian source field under a time-weighted budget constraint on actuation costs. The plasma model is then coupled to CNN-based soft actor-critic and twin-delayed deterministic policy-gradient agents through a GPU-native JAX solver that is highly optimized for fast online training. In the turbulence-suppression task, the learned budget-aware schedule achieves the lowest time-integrated turbulent flux for a given consumed budget, outperforming both constant and linearly decreasing baselines across all tested unseen initial conditions. In the inverse zonal-break task, the agent discovers an up-down antisymmetric actuation pattern that induces radial $E\times B$ convection, disrupts the zonal structure, and sustains the turbulent state. A physics-informed warm-buffer initialization facilitates this discovery, as random exploration alone struggles to locate the narrow optimal manifold within the vast action space. These results demonstrate reinforcement learning as a practical trajectory optimizer for nonlinear plasma dynamics, such as turbulence control, and highlight the importance of physics guidance in such applications.

Comments28 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑