AI 中文总结
该综述梳理了2025-2026年神经网络优化的进展,沿四个轴组织相关方法,指出矩阵感知方法是进步但无AdamW的通用替代方案,需结合多因素评估优化器。
AI 中文摘要
2025-2026年的神经网络优化已不再能用一系列新的Adam变体来充分描述。设计空间已从坐标扩展到矩阵和层,从固定训练范围扩展到随时间变化的策略,从数学更新规则扩展到必须能在分片和低精度计算中保留的状态表示。本综述沿四个基本独立的轴组织了近期的优化器和训练优化方法:时间估计、更新几何、范围管理以及表示与系统。它关联了Muon的谱归一化、Shampoo和SOAP的历史矩阵统计、自适应与混合矩阵方法、内存高效优化器、无调度训练、小批量校正以及量化优化器状态。核心经验结论刻意避免了绝对化:感知矩阵的方法确实是一项进步,但不存在能替代AdamW的通用方案。排名会随模型规模、数据与参数的比率、批量大小、调度、参数划分、调优预算,以及目标指标是token、FLOPs、墙钟时间还是内存而变化。实际结果是优化器设计的组合视角,以及评估优化器主张的更严格协议。
英文摘要
Neural-network optimization in 2025-2026 is no longer well described as a succession of new Adam variants. The design space has expanded from coordinates to matrices and layers, from fixed training horizons to policies over time, and from mathematical update rules to state representations that must survive sharding and low-precision computation. This survey organizes recent optimizers and training optimization methods along four largely independent axes: temporal estimation, update geometry, horizon management, and representation and systems. It connects the spectral normalization of Muon, the historical matrix statistics of Shampoo and SOAP, adaptive and hybrid matrix methods, memory-efficient optimizers, schedule-free training, small-batch corrections, and quantized optimizer states. The central empirical conclusion is deliberately non-triumphal: matrix-aware methods represent a genuine advance, but there is no context-independent replacement for AdamW. Rankings change with model scale, data-to-parameter ratio, batch size, schedule, parameter partition, tuning budget, and whether the target metric is tokens, FLOPs, wall-clock time, or memory. The practical consequence is a compositional view of optimizer design and a stricter protocol for evaluating optimizer claims.