PEAT:DNN训练中GPU内核验证的伪误差评估
PEAT: Pseudo-Error Assessment for GPU Kernel Validation in DNN Training
浏览论文内容
中文总结 AI 辅助
针对DNN训练中内核验证耗时且需大量存储的问题,提出PEAT轻量级框架,通过回放和频率故障注入收集状态,分析误差特征,提供充分条件关联误差模型,并在NVIDIA和AMD GPU上验证。
中文摘要 AI 辅助
深度神经网络(DNN)在各个领域被广泛采用,推动了与DNN训练系统相关的软件栈开发的新兴趋势。例如,许多代码已被移植到不同的框架中,或为利用GPU或特定领域加速器的计算能力而开发。然而,在DNN训练中验证内核实现既耗时又通常需要大量存储。具体而言,这提出了一个基本问题:当新实现被集成到DNN训练流程中时,如何表征其行为。不幸的是,据我们所知,这一问题在文献中尚未得到充分研究。为解决这一不足,我们提出了PEAT——一个用于DNN训练中GPU内核验证的轻量级检查框架,用于伪误差评估。首先,受传统故障注入(FI)启发,PEAT的Profiler在训练流程中调用逐操作内核,以收集DNN模型的状态(如检查点和激活值)。更重要的是,Profiler引入了两种简单而有效的技术:回放故障注入和基于频率的运行时故障注入,利用训练过程中的持久内核调用。其次,PEAT的Analyzer对剖析的误差进行表征,揭示内核与黄金内核相比误差分布中的一些特征。最后,PEAT的Detector提供了一些作为充分条件的指导方针,使若干已知误差模型与特征模式相关联。我们通过使用来自两大最流行供应商NVIDIA V100和AMD MI250的GPU,在从视觉任务到语言模型的多种AI模型上,针对预训练和微调场景,展示结果和分析,证明了我们方法的适用性。
英文摘要
Deep neural networks (DNNs) are widely adopted in various fields, driving an emerging trend in developing software stacks associated with DNN training systems. For example, many codes have been ported across different frameworks or developed to leverage the computing power of GPUs or domain-specific accelerators. However, validating a kernel implementation in DNN training is time-consuming and generally requires massive storage. Specifically, this poses a fundamental question: how to characterize the behavior of a new implementation when it is integrated into a DNN training flow. Unfortunately, this problem is not well investigated in the literature, to the best of our knowledge. To address this shortcoming, we present PEAT - a lightweight inspection framework for \underline{P}seudo-\underline{E}rror \underline{A}ssessment associated with GPU kernel validation in DNN \underline{T}raining. Firstly, inspired by conventional fault injection (FI), PEAT's Profiler invokes an operation-wise kernel in a training flow to collect a DNN model's states (e.g., checkpoints and activations). More importantly, the Profiler introduces two simple yet effective techniques, playback FI and frequency-based runtime FI, leveraging persistent kernel calling during the training process. Secondly, PEAT's Analyzer characterizes profiled errors, revealing some signatures from the error distribution of a kernel compared to the golden one. Lastly, PEAT's Detector provides some guidelines as a sufficient condition, which enables associating several well-known error models with signature patterns. We demonstrate the applicability of our approach by presenting the results and analysis using GPUs from the two most popular vendors, NVIDIA V100 and AMD MI250, on various AI models, from vision tasks to language models, for both pretraining and finetuning scenarios.
发表机构
- Seoul National University(首尔大学)
- Van-Lang Institute of Semiconductor Technology (VIST)(文朗半导体技术研究所)
- Vietnam National University(越南国立大学)
- Moreh Vietnam(Moreh越南)
机构由 AI 辅助整理,请以论文原文为准。