发表机构
MIT(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明大型语言模型可训练以容忍硬件不可靠性,且规模增大时错误韧性增强,基于修正缩放定律推测其可能形式容错,为低能耗硬件上的AI推理提供节能路径。
AI 中文摘要
新兴计算机硬件往往以可靠性换取能源效率;我们在此表明,大型语言模型(LLMs)可以被训练以容忍这种不可靠性,并且随着模型规模的增大,其错误韧性实际上会增强而非退化。从在模拟故障数字硬件上进行的40,000 GPU小时的训练运行中推断出的修正神经缩放定律量化了这一趋势,并表明模型学会在“良好”的纠错码内进行计算,无论模型变得多大,这些纠错码的相对开销都保持有限。这一发现使我们推测,经过适当训练的LLMs可能在形式上具有容错性;如果属实,在低能耗、有故障的硬件上运行AI推理可能是相对于现状实现大幅能源节省的一条途径。
英文摘要
Emerging computer hardware often trades reliability for energy efficiency; here we show that large-language models (LLMs) can be trained to tolerate this unreliability, and that rather than degrading, their error resilience actually increases as they grow. Modified neural scaling laws inferred from 40,000 GPU-hours of training runs on simulated faulty digital hardware quantify this trend and suggest that models learn to compute within "good" error-correcting codes, whose relative overhead remains finite no matter how large the model gets. This finding leads us to conjecture that appropriately trained LLMs may be formally fault-tolerant; if true, running AI inference on low energy, faulty hardware may be a path to substantial energy savings over the status quo.