arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27578physics.plasm-phcs.LG

面向科学基础模型的大规模异构数据组织:以核聚变研究为例

Towards Large-Scale Heterogeneous Data Organization for Scientific Foundation Models: A Nuclear Fusion Case Study

Nathaniel Chen, Kouroche Bouchiat, Peter Steiner, Azarakhsh Jalalvand, SangKyeun Kim, Egemen Kolemen

首次发表
浏览论文内容

中文总结 AI 辅助

本文以核聚变研究为例,针对科学基础模型训练所需的大规模异构数据问题,表征了相关数据特性并分析了复杂度,提出了多模态波动数据表征模板,为相关领域提供了参考。

中文摘要 AI 辅助

训练有效的基础模型需要海量且有序的数据集,但核聚变等科学领域因数据高度异构且稀疏而面临独特挑战。本文对开发此类模型所用数据进行了表征:包含超过20种传感器类型,采样率跨度达5个数量级,混合张量结构(点测量、频谱图、图像)以及非平稳物理特性。我们分析了输入复杂度,并讨论了时间上下文与频率分辨率之间的权衡。该分析为大规模多模态波动数据的表征提供了模板,对多模态控制系统和核聚变研究均具有意义。

英文摘要

Training effective foundation models requires massive and organized datasets, yet scientific domains such as nuclear fusion present unique challenges due to largely heterogeneous and sparse data. Here we characterize the data used in developing such a model: with over 20 sensor types spanning 5 orders of magnitude in sampling rate, mixed tensor structures (point measurements, spectrograms, images), and nonstationary physics. We analyze our input complexity and discuss trade-offs between temporal context and frequency resolution. Our analysis provides a template for representing multi-modal fluctuation data at scale, with implications for both multi-modal control systems and nuclear fusion.

补充信息

↑