发表机构
University of Bologna; CINECA(博洛尼亚大学; 意大利国家科研计算中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对数据中心计算节点建模需求,提出SeT-Diff基础模型,采用基于扩散的方法,在真实超级计算机数据集实验中,其重建任务平均绝对误差为0.0470,有零样本排列稳定性,单个模型可有效执行多种任务,是高性能计算系统有效的数据驱动数字孪生。
AI 中文摘要
数据中心及其计算节点需要能够对工作负载、环境参数和物理指标的复杂相互作用进行建模的准确且灵活的数字孪生。当前用于高性能计算及其遥测的机器学习方法通常依赖于为单个任务量身定制的匿名、固定位置传感器变量的静态子集。我们提出了SeT-Diff,首个用于计算节点遥测和时间序列的基础模型。与刚性架构不同,我们基于扩散的方法在每个传感器的语义描述上对生成过程进行条件设定,将系统动态与数据集结构解耦。在真实世界超级计算机数据集上的实验表明,重建任务的平均绝对误差为0.0470,SeT-Diff具有零样本排列稳定性,单个预训练模型能有效执行数据插补、预测和虚拟传感,是高性能计算系统有效的数据驱动数字孪生。
英文摘要
Data centers and their compute nodes require accurate and flexible digital twins capable of modeling the complex interplay of workloads, environmental parameters, and physical metrics. Current machine learning approaches for HPC and its telemetry typically rely on a static subset of anonymous, fixed-position sensor variables tailored to single tasks. Consequently, these models become obsolete when target tasks change or sensor metrics vary. We propose SeT-Diff, the first foundational model for compute node telemetry and time-series. Unlike rigid architectures, our diffusion-based approach conditions the generative process on each sensor's semantic description, decoupling the system dynamics from the structure of the dataset. Experiments on a real-world supercomputer dataset demonstrate a Mean Absolute Error (MAE) of 0.0470 on reconstruction tasks. SeT-Diff exhibits zero-shot permutation stability, maintaining accuracy with negligible degradation even when sensors are shuffled. A single pre-trained model effectively performs data imputation, forecasting, and virtual sensing - achieving a 0.033 MAE in thermal inference - making SeT-Diff an effective data-driven digital twin for HPC systems.