arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

让模拟训练规模化:协同设计映射、优化器与转换器

Making Analog Training Scale: Co-Designing Mapping, Optimizer, and Converters

Zhaoxian Wu, Tayfun Gokmen, Omobayode Fagbohungbe, T. Patrick Xiao, Tianyi Chen

arXiv 2609.36584首次发表:更新:

发表机构

Cornell Tech and Cornell University; IBM T. J. Watson Research Center; Sandia National Laboratories(康奈尔科技学院和康奈尔大学; IBM T. J. Watson 研究中心; 桑迪亚国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对模拟存内计算训练深度模型的硬件非理想性挑战,提出系统-算法协同设计,通过映射、优化器与转换器协同优化,实现1.23亿参数Transformer训练,验证损失缩放接近数字训练。

AI 中文摘要

模拟存内计算(AIMC)通过在存储权重的位置直接执行矩阵运算,为模型训练提供了一种替代方案。然而,由于严重的硬件非理想性,包括具有有限动态范围和写入粒度的物理权重、具有有限分辨率的模数转换器,以及有噪声且不对称的更新,将AIMC扩展到训练现代深度模型仍然是一个开放的挑战。基于梯度累积对精度和舍入误差敏感的洞察,我们采用了一种混合精度训练范式:在模拟域中执行前向和反向矩阵乘法,同时在数字域中计算权重梯度。为了实现可扩展的训练,我们提出了一种整体性的系统-算法协同设计,该设计协同优化权重映射以确保物理和逻辑权重分布的良好条件,将预条件优化器与阈值触发的开环脉冲相结合以稳定训练轨迹,并调整转换器动态范围以抑制量化误差。通过使用电化学RAM测量校准的硬件校准架构仿真进行评估,我们的框架将Transformer训练扩展到高达1.23亿个参数,验证损失缩放为L∝N^{-0.231},其中N为参数数量,与数字训练的L∝N^{-0.238}相当。

英文摘要

Analog in-memory computing (AIMC) offers an alternative for model training by executing matrix operations directly where weights are stored. However, scaling AIMC to train modern deep models remains an open challenge due to severe hardware non-idealities, including physical weights with finite dynamic range and write granularity, analog-digital converters with finite resolution, and noisy and asymmetric updates. Guided by the insight that gradient accumulation is sensitive to precision and rounding errors, we adopt a mixed-precision training paradigm: executing forward and backward matrix multiplications in the analog domain while computing weight gradients in the digital domain. To enable scalable training, we present a holistic system-algorithm co-design that co-optimizes weight mapping to ensure well-conditioned physical and logical weight profiles, couples a preconditioned optimizer with threshold-triggered open-loop pulsing to stabilize training trajectories, and aligns converter dynamic ranges to suppress quantization errors. Evaluated via hardware-calibrated architectural simulations calibrated with electrochemical RAM measurements, our framework scales Transformer training up to $123\text{M}$ parameters with validation loss scaling as $L\propto N^{-0.231}$, where $N$ is the parameter count, comparable to $L\propto N^{-0.238}$ for digital training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑