发表机构
Thomas Jefferson National Accelerator Facility; Pacific Northwest National Lab(托马斯·杰斐逊国家加速器实验室; 太平洋西北国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种结合代理建模与强化学习的数据驱动框架,用于优化核物理散射实验中的靶极化,利用高斯过程模型预测极化并训练智能体,使操作效果提升近2倍。
AI 中文摘要
在核物理实验中,动态极化靶的操作依赖于对微波频率的连续调节,以补偿辐射损伤和材料性质的变化,这一任务传统上由专家操作员通过手动试错完成。本工作提出了一种数据驱动的控制框架,将代理建模与强化学习相结合,以优化靶极化。利用APOLLO低温靶系统的运行数据,我们训练并评估了多层感知机和高斯过程回归模型,用于预测极化随微波频率、束流强度和累积辐射剂量的变化。我们表明,基于高斯过程的模型提供了校准的不确定性估计,并能可靠地识别训练分布之外的区域,而多层感知机对分布偏移的敏感性有限。为了在多个靶样本之间实现学习和控制,我们引入了一种高斯过程近似,并将代理模型嵌入标准化的模拟环境中。使用下置信界奖励公式训练强化学习智能体,该公式在性能最大化和不确定性之间取得平衡。我们能够证明,利用我们的强化学习智能体,操作员的行动效果几乎提高了2倍。
英文摘要
The operation of dynamically polarized targets in nuclear physics experiments relies on continuous tuning of the microwave frequency to compensate for radiation damage and evolving material properties, a task that is traditionally performed through manual trial-and-error by expert operators. This work presents a data-driven control framework that combines surrogate modeling with reinforcement learning to optimize the target polarization. Using operational data from the APOLLO cryogenic target system, we train and evaluate multilayer perceptron and Gaussian process regression models to predict polarization as a function of microwave frequency, beam current, and accumulated radiation dose. We show that Gaussian process-based models provide calibrated uncertainty estimates and reliably identify regions outside the training distribution, while MLPs exhibit limited sensitivity to distributional shift. To enable learning and control across multiple target samples, we introduce a Gaussian process approximation and embed the surrogate model within a standardized simulation environment. A reinforcement learning agent is trained using a lower-confidence-bound reward formulation that balances performance maximization against uncertainty. We are able to show an almost 2x improvement on the operators actions utilizing our RL agent.