AI 中文总结
该研究针对内存高效优化器的状态异质性提出自适应对数空间量化,在AdamW、CAME等模型的语言模型预训练实验中,可显著降低优化器存储并保持良好困惑度表现。
AI 中文摘要
低精度优化器状态方法通常针对密集Adam式矩设计,但内存高效优化器维护因式分解、基于置信度或投影的状态,其量化误差传播方式不同。我们在语言模型预训练的优化器状态轨迹中表征这种异质性,提出自适应对数空间(Adaptive Log-Space, AL)量化,这是一种针对非负状态的块级表示,每个块自适应其非零范围同时保留精确零值。AL8和AL16与独立的符号动量编码及特定状态的精度选择相结合。在总计214.7 GPU小时的96次运行中,我们评估AdamW、Adafactor、CAME和APOLLO路径。在2万步的TinyLlama-1.1B基准测试中,采用AL8二阶矩和8位均匀动量的AdamW达到72.90困惑度,而FP32的困惑度为72.48,8位动态量化基线的困惑度为73.54,同时优化器状态存储从8392.7 MiB降至2119.2 MiB。CAME对其非负状态需要更高精度:AL16达到86.16困惑度,FP32为86.68,全AL8则为90.19。在10万步的GPT-2实验中,拓扑感知参数保护将量化Adafactor的后期损失差距从+0.1185降至+0.0159。这些结果支持针对状态和拓扑感知的优化器量化,端到端比较使用单一训练种子并作为经验测量报告。
英文摘要
Optimizer-state quantization is commonly designed for Adam's dense, parameter-aligned first- and second-moment arrays. This abstraction breaks for memory-efficient optimizers, whose states may be factored, confidence-modulated, or maintained in a projected space, so similar reconstruction error can produce different update error. We formulate optimizer-state quantization as a joint problem over representation, topology, and update semantics. We then introduce Adaptive Log-Space (AL) quantization for non-negative states. AL fits each block's observed nonzero logarithmic interval and reserves a separate code for exact zero, enforcing $q = 0 \Leftrightarrow x = 0$; signed momentum and state precision remain independently selectable. Controlled probes show that adaptive ranges reduce update error and temporal drift, exact-zero reservation preserves dormant states, and state topology constrains useful block granularity. End-to-end language-model training evaluates the resulting policy across dense, factored, confidence, and projected optimizer states. On TinyLlama-1.1B, AL8 with uniform 8-bit momentum reaches 72.90 perplexity versus 73.54 for bitsandbytes 8-bit AdamW, with comparable optimizer-state storage and higher throughput. CAME matches reference-level final perplexity across three seeds when its non-negative states use AL16, while a semantic grouping-and-protection policy closes most of quantized Adafactor's 100K-step late-loss gap. These results make state topology and update semantics first-class design constraints for optimizer quantization.
Comments17 pages, 5 figures, 7 tables. Substantially revised framing and evidence; added a matched adaptive-log comparison, multi-seed validation, and the completed Adafactor block-size/grouping grid. Updated reproducibility mapping and evaluation scope. Code: https://github.com/yanfeiwong/adafactor-8bit. Artifacts: https://github.com/yanfeiwong/al-quantization