发表机构
University of Oxford; University of Chinese Academy of Sciences; Shanghai Innovation Institute; Westlake University(牛津大学; 中国科学院大学; 上海创新研究院; 西湖大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对托卡马克实时平衡重建中实验数据稀缺的问题,提出基于物理模拟的合成预训练加真实数据微调方法,在仅用1%真实数据时显著降低重建误差,提升数据稀缺下的重建精度。
AI 中文摘要
实时平衡重建对于托卡马克等离子体控制至关重要。数据驱动的替代模型通常在高品质标记实验数据上训练,但产生足够用于替代学习的大量标记数据需要许多昂贵的放电炮次,而在新活动开始或新装置运行时几乎没有此类数据。为降低这一成本,我们提出在由基于物理的模拟低成本生成的模拟平衡上进行预训练,然后在有限的真实数据上进行微调。我们在三种难度递增的设置下进行评估,即分布内、分布外(形状)和跨活动(时间)分割,这些设置反映了重建实际部署的方式。实验结果验证了当真实标签稀缺时,合成预训练高度有效:仅在全可用真实数据集的1%上进行微调,与在相同数量真实数据上从头训练的模型相比,在分布外形状下重建极向磁通量的每样本归一化均方根误差(nRMSE)降低了约58%,在跨活动外推下降低了64%。这些现实评估表明,合成预训练提高了真实托卡马克活动面临的数据稀缺状态下的平衡重建精度。
英文摘要
Real-time equilibrium reconstruction is essential for tokamak plasma control. Data-driven surrogates are usually trained on high-quality labeled experimental data, but producing a large amount of labeled data sufficient for surrogate learning requires many expensive discharge shots, and few exist at the start of a new campaign or the operation of a new device. To reduce this cost, we propose to pre-train on simulated equilibria, generated at low cost by physics-based simulation, then fine-tune on limited real data. We evaluate under three settings with increasing difficulties, i.e., in-distribution, out-of-distribution (shape), and cross-campaign (temporal) splits, which mirror how reconstruction is actually deployed. Experimental results validate that synthetic pre-training is highly effective when few real labels are available: fine-tuning on only 1% of the full available real dataset cuts the per-sample normalized root-mean-square error (nRMSE) of the reconstructed poloidal flux by roughly 58% under out-of-distribution shapes and 64% under cross-campaign extrapolation, compared to a model trained from scratch on the same amount of real data. These realistic evaluations show that synthetic pre-training improves equilibrium reconstruction accuracy in the data-scarce regime that real tokamak campaigns face.