AI 中文总结
本研究提出p-自旋玻璃网络架构,通过三元量化等技术实现内存与样本效率,支持单批次稳定收敛,消除深度学习对大批次的需求,为持续学习和边缘AI奠定基础。
AI 中文摘要
现代序列模型严重依赖庞大的内存占用和大批次随机优化,这些障碍限制了样本效率和持续学习能力。我们提出了p-自旋玻璃网络(p-Spin Glass Network),一种新型架构,可克服这些限制,从结构上管控优化方差,并具备四项显著能力:1. 它实现内存高效:原生三元量化将内部参数压缩8倍,而精确的隐式梯度严格将激活内存限制为O(B·T·D);2. 它展现样本效率,在使用8倍更少训练序列的同时,达到Transformer基线的渐近性能;3. 该方法支持单批次稳定性,且在随机微批次大小为1时实现平滑、单调收敛;4. 最后,这种稳定性具备模态无关性,在离散子词和长时序未压缩原始字节流上均保持稳健的时间信用分配。最终,本研究消除了稳定深度学习对大批次的需求,为持续学习和边缘AI奠定了基础。
英文摘要
Modern sequence models heavily rely on massive memory footprints and large-batch stochastic optimization, barriers that restrict sample efficiency and continual learning. We introduce the $p$-Spin Glass Network, a novel architecture that overcomes these limitations, structurally manages optimization variance and yields four noticeable capabilities: 1. It enforces memory efficiency: native ternary quantization compresses internal parameters by $8\times$, while exact implicit gradients strictly bound activation memory to $\mathcal{O}(B \cdot T \cdot D)$. 2. it demonstrates sample efficiency, matching the asymptotic performance of a Transformer baseline while utilizing $8\times$ fewer training sequences. 3. Method enables single-batch stability and smooth, monotonic convergence at a stochastic micro-batch size of $1$. 4. Finally, this stability proves modality-agnostic, maintaining robust temporal credit assignment across both discrete subword and long horizon uncompressed raw byte streams. Ultimately, this work removes large batch requirement for stable deep learning, establishing a foundation for continuous learning and edge AI.