可用性感知训练何时值得?可预测计算调度下中断弹性优化的基准与实证研究
When Is Availability-Aware Training Worth It? A Benchmark and Empirical Study of Interruption-Resilient Optimization Under Predictable Compute Schedules
- RotaStellar
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过基准OrbitTrace和强基线对比,发现当优化器状态可保留时,可用性间隙几乎无代价,专门的中断弹性优化仅在特定大模型频繁短中断场景下略有帮助。
AI中文摘要:
在非平稳但可预测的计算可用性下进行训练(如卫星处于日食期、占空比受限的边缘设备、功率受限的数据中心)通常被认为需要专门的、可用性感知的优化器。我们检验了这一前提。我们发布了OrbitTrace,一个包含50条基于物理的可用性轨迹的基准,这些轨迹来自对三个轨道体制下实时双行元素集的SGP4传播,并提出了一个可证伪的问题:当可用性间隙中断训练时,专门的恢复策略是否值得,还是胜任的检查点-恢复就足够了?我们的核心发现:当优化器状态可以在间隙期间保留时,该间隙基本上是免费的。一个强大的检查点基线,能够恢复完整的优化器状态并在有效(活动)时间内索引其学习率调度,在CIFAR-10/ResNet-18上与不间断训练的差异仅在数据排序噪声范围内,在GPT-2/AdamW任务上则完全一致。先前报道的可用性感知方法的优势,包括我们自己的三支柱方法AAT,几乎完全源于与一个在墙钟时间上索引其调度的弱基线进行比较。相对于强基线,反应式适应在平稳、优化器状态丢失和分布漂移机制中均无优势。我们隔离了一个狭窄的机制,其中它有所帮助:对于优化器状态无法跨间隙持久化且被频繁、短暂暂停中断的大模型,重建衰减的优化器矩平均仅恢复状态丢失惩罚的约21%(且在不同种子间不稳健);对于日食尺度的间隙,这种优势消失,因为衰减的矩与零无法区分。我们的贡献包括一个基准、一个强大的可复现基线协议,以及关于中断弹性优化何时值得其复杂性、何时不值得的清晰刻画。
英文摘要:
Training under non-stationary but predictable compute availability (satellites under eclipse, duty-cycled edge devices, power-capped datacenters) is often framed as needing specialized, availability-aware optimizers. We test that premise. We release OrbitTrace, a benchmark of 50 physics-grounded availability traces from SGP4 propagation of live two-line element sets across three orbital regimes, and ask a falsifiable question: when an availability gap interrupts training, is a specialized resumption strategy worth it, or is competent checkpoint-and-resume enough? Our central finding: when optimizer state can be preserved across a gap, the gap is essentially free. A strong checkpoint baseline that restores full optimizer state and indexes its learning-rate schedule in effective (active) time matches uninterrupted training to within data-ordering noise on CIFAR-10/ResNet-18, and exactly on a GPT-2/AdamW task. Advantages previously reported for availability-aware methods, including our own three-pillar method AAT, arise almost entirely from comparison against a weak baseline that indexes its schedule on wall-clock time. Against the strong baseline, reactive adaptation provides no advantage across stationary, optimizer-state-loss, and distribution-drift regimes. We isolate one narrow regime where it helps: for large models whose optimizer state cannot be persisted across gaps and that are interrupted by frequent, short pauses, reconstructing a decayed optimizer moment recovers only ~21% of the state-loss penalty on average (and not robustly across seeds); this vanishes for eclipse-scale gaps, where the decayed moment is indistinguishable from zero. Our contributions are a benchmark, a strong reproducible baseline protocol, and a clear characterization of when interruption-resilient optimization is worth its complexity, and when it is not.