AI 中文总结
研究针对自适应优化器在不同训练阶段更新规则相同的问题,提出PsiLogic优化器,通过双指数移动平均控制主动抵消项,经FairBench基准测试,在多个领域表现优异,还发布开源实现等支持验证。
AI 中文摘要
自适应优化器如Adam和AdamW,无论训练处于混沌早期阶段还是接近收敛,都应用相同的更新规则。我们引入了PsiLogic,一种通过由尺度归一化梯度范数的双指数移动平均(EMA)控制的动态主动抵消项增强Adam的优化器。由此产生的混沌检测器在梯度统计不稳定时增强阻尼,并在训练稳定时逐渐消失为零,提供了无需手动调整时间表的隐式热身。我们使用FairBench对PsiLogic与Adam、AdamW和Lion进行评估,这是一种具有每个优化器学习率扫描、每个种子相同初始化以及Welch t检验的可重复基准测试协议。在NVIDIA H100 80GB参考运行中(4个领域,3个种子,2000步,bf16 AMP),PsiLogic在四个领域中的三个中实现了最佳验证指标:NLP困惑度7.79±0.18对比8.17±0.08(AdamW,p = 0.049),ViT top-1准确率0.244±0.006对比0.223±0.002(AdamW,p = 0.015),以及ResNet top-1准确率0.222±0.001对比0.172±0.004(Adam,p = 0.001)。在扩散方面,验证MSE与Adam/AdamW在统计上相当(p = 0.49)。ResNet准确率与AdamW在三个种子上是数值上的平局且无显著性差异(p = 0.44)。峰值GPU内存跨优化器可比;PsiLogic在变压器密集型领域产生1.2 - 1.8倍的时钟开销(受实现限制)。我们发布了开源的PyTorch实现、完整的FairBench工具包以及所有原始CSV输出以支持独立验证。
英文摘要
Adaptive optimizers such as Adam and AdamW apply the same update rule regardless of whether training is in a chaotic early phase or near convergence. We introduce PsiLogic, an optimizer that augments Adam with a dynamic Active Cancellation Term gated by a dual exponential moving average (EMA) of scale-normalized gradient norms. The resulting chaos detector strengthens damping when gradient statistics are unstable and fades to zero as training stabilizes, providing an implicit warmup without a hand-tuned schedule. We evaluate PsiLogic against Adam, AdamW, and Lion using FairBench -- a reproducible benchmark protocol with per-optimizer learning-rate sweeps, identical initialization per seed, and Welch t-tests. On an NVIDIA H100 80GB reference run (4 arenas, 3 seeds, 2000 steps, bf16 AMP), PsiLogic achieves the best validation metric in three of four arenas: NLP perplexity 7.79 +/- 0.18 vs. 8.17 +/- 0.08 (AdamW, p = 0.049), ViT top-1 accuracy 0.244 +/- 0.006 vs. 0.223 +/- 0.002 (AdamW, p = 0.015), and ResNet top-1 accuracy 0.222 +/- 0.001 vs. 0.172 +/- 0.004 (Adam, p = 0.001). On diffusion, validation MSE is statistically tied with Adam/AdamW (p = 0.49). ResNet accuracy vs. AdamW is a numerical tie without significance at three seeds (p = 0.44). Peak GPU memory is comparable across optimizers; PsiLogic incurs 1.2--1.8x wall-clock overhead on transformer-heavy arenas (implementation-bound). We release an open-source PyTorch implementation, the full FairBench harness, and all raw CSV outputs to support independent verification.
Comments9 pages, 4 figures; code at https://github.com/Troxter222/psilogic