应用故障与机器计算效率
Application Failures and Machine Computational Efficiency
AI总结:
该研究针对应用故障率较高的情况,提出了基于使用情况而非时间评估百亿亿次级科学计算机正常运行时间效率的框架,结合Frontier超算数据验证了方法,可优化检查点间隔并提升计算资源利用率。
AI中文摘要:
当应用故障率较高时,我们提出了一种用于评估百亿亿次级科学计算机正常运行时间效率的框架,这是当前顶级科学计算平台和大型AI训练装置面临的情况。科学计算平台的特点是其应用的异构性,我们认为这种多样性要求故障率和平均故障间隔时间应基于使用情况(例如节点小时)而非当前惯用的时间来指定。我们考虑了故障、检查点和重启成本导致的使用损失项,并更新了Daly(2006)的框架,允许用户指定最优检查点使用间隔以最小化此类损失。我们推导了机器计算效率,该效率指定了可用于科学计算的预期资源分配比例。我们使用橡树岭国家实验室Frontier超级计算机一年的生产运行时间数据说明了该方法。
英文摘要:
We present a framework for evaluating uptime efficiency of Exascale-class scientific computers when application failure rates are appreciable. This is the situation that confronts current leadership-class scientific computing platforms and large AI training installations. What distinguishes scientific computing platforms is the heterogeneity of their applications. We argue that this diversity requires that failure rates and mean intervals between failures should be specified in terms of \emph{usage} (e.g. node-hours) rather than time, as is currently customary. We consider the usage loss terms due to failures, to checkpointing, and to restart costs, and update the framework of Daly (2006) allowing users to specify optimal checkpointing usage intervals that minimize such losses. We derive the machine computational efficiency, which specifies the expected fractional resource allocation that is available for scientific computation. We illustrate the methodology using one year of production runtime data from the \emph{Frontier} supercomputer at Oak Ridge National Laboratory.