在NVIDIA B300上运行多节点全微调:基于遥测的分类、负面结果与操作强化的现场报告
Operating Multi-Node Full Fine-Tuning on NVIDIA B300: A Field Report on Telemetry-Based Triage, Negative Results, and Operational Hardening
- Samsung SDS(三星SDS)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究分享了在NVIDIA B300上微调327.6亿参数模型的操作经验,提供功耗分类表、负面结果、校准扩展数据及故障解决方案,强调操作层面的实践要点。
AI中文摘要:
我们报告了在16块NVIDIA B300(两个节点,采用FSDP / ZeRO-3)上对参数规模为327.6亿的稠密模型(Qwen3-32B)进行全微调的操作经验——这是首批公开的关于该加速器的现场报告之一。我们不提出新算法,所使用的各个机制均为成熟实践;我们的贡献在于整合了现场经验以及针对新硬件的一组校准测量值。具体而言,我们提供了四类从业者可用的成果:(1)一张经B300校准的功耗分类表,可通过板卡瓦数区分计算/通信/数据饥饿/检查点或死锁/空闲状态(当NCCL挂起时,利用率百分比读数为100%);(2)一组真实的负面结果,打破了该规模下常见的优化误区:一项对照A/B测试显示,每步NFS读取速度与预分词本地缓存(约5.3万词元/秒)相当,原因是语料库适配页缓存且作业受计算限制;以及将早期的“吞吐量崩溃”归因于NFS/CPU争用而非存储介质限制;(3)B300上校准的4/8/16-GPU强扩展与GPU小时数(近线性,符合该场景预期;我们报告绝对值作为参考数据);(4)一个实际故障案例——来自按秩词元打包失衡的周期结束NCCL死锁,以及一个耗时2.7秒的预运行不变量检查门和一个外部监视器,可将数小时的静默故障转为即时弃权(不执行)。该死锁及其补救措施对应PyTorch已记录的Join/均等化至最小值实践;我们将我们的实现与现有技术进行对比,并报告故障造成的GPU小时数及该检查门节省的GPU小时数。可迁移的核心结论是操作层面而非算法层面:对于依赖数据的数据并行作业,应监控功耗而非利用率,且在启动前验证不变量——通过冒烟测试并不代表完整运行是安全的。
英文摘要:
We report operational experience full-fine-tuning a 32.76B-parameter dense model (Qwen3-32B) on 16 x NVIDIA B300 (two nodes, FSDP / ZeRO-3) -- among the first published field accounts on this accelerator. We claim no new algorithm. The individual mechanisms we use are established practice; our contribution is the integrated field experience and a set of calibrated measurements on new hardware. Concretely we offer four practitioner artifacts. (1) A B300-calibrated power-draw triage table that distinguishes compute / communication / data-starvation / checkpoint-or-deadlock / idle by board wattage (utilization% reads 100% during an NCCL hang). (2) A set of honest negative results that dispel common optimization folklore at this scale: a controlled A/B in which per-step NFS reading matches a pretokenized local cache (~53k tok/s) because the corpus fits in page cache and the job is compute-bound; and a reconstruction of an earlier "throughput collapse" as NFS/CPU contention rather than a storage-medium limit. (3) Calibrated 4/8/16-GPU strong-scaling and GPU-hour numbers on B300 (near-linear, as expected in this regime; we report absolute values as reference data). (4) A worked failure case -- an epoch-end NCCL deadlock from per-rank token-packing imbalance -- together with a 2.7-second pre-run invariant gate and an external watcher that turn multi-hour silent failures into instant rejections. This deadlock and its remedy correspond to PyTorch's documented Join / equalize-to-minimum practice; we position our instantiation against that prior art and report the GPU-hours the failure cost and the gate saves. The transferable takeaway is operational, not algorithmic: for data-dependent data-parallel jobs, watch power rather than utilization, and verify invariants before launch -- a passing smoke test is not evidence of a safe full run.