arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04266cs.ARcs.LG

存算一体注意力:具有RC可调温度的时域模拟Softmax电路

Compute-in-Memory Attention: A Time-Domain Analog Softmax Circuit with RC-Tunable Temperature

  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

Ankur Singh, Ashish Gautam, Shruti R. Kulkarni, Guojing Cong

AI总结:

该研究提出一种GlobalFoundries 22nm FDSOI工艺的时域模拟Softmax电路,无需模数转换,通过RC时间常数实现温度可调,128元件架构在低功耗下误差小,用于Transformer模型验证损失接近理想基线。

AI中文摘要:

Softmax是Transformer注意力机制中的关键操作,但其指数运算和归一化会给存算一体(CIM)加速器带来显著开销,尤其是当模拟注意力分数必须先转换到数字域时。本研究采用GlobalFoundries 22纳米全耗尽绝缘体上硅(FDSOI)工艺,提出一种可调温度的模拟Softmax电路,可直接处理CIM生成的分数电压,无需中间模数转换。每个输入分数通过共享下降斜坡转换为时域事件,对应的比较器跳变对RC衰减参考进行采样以生成指数权重,随后由片内归一化级处理。与依赖晶体管弱反相行为实现指数运算的模拟Softmax电路不同,所提架构通过斜坡斜率和RC时间常数控制Softmax响应,支持可编程有效温度。该128元件架构通过晶体管级和布局后提取仿真进行评估,涵盖多电平输入向量、电容变化与失配、工艺与温度变化、蒙特卡洛分析及共享互连寄生参数。完整的128元件实现(含共享全局斜坡电路)占用9453.42平方微米,每个复制的Softmax元件占用70.2平方微米。该电路在13.44毫瓦总功率下实现242.97纳秒的评估延迟,对应每个输出元件25.5皮焦。同步128元件评估相对于理想Softmax响应的均方根误差(RMSE)为24.46毫伏。提取的电路特性进一步被整合到基于MemTorch的硬件感知Transformer模型中,其中所提Softmax的验证损失在理想Softmax基线的2.5%以内。

英文摘要:

Softmax is a key operation in Transformer attention, but its exponentiation and normalization add significant overhead in compute-in-memory (CIM) accelerators, especially when analog attention scores must first be converted to the digital domain. This work presents a tunable-temperature analog softmax circuit in GlobalFoundries 22-nm fully depleted silicon-on-insulator (FDSOI) technology that operates directly on CIM-generated score voltages without intermediate analog-to-digital conversion. Each input score is converted into a time-domain event using a shared falling ramp. The corresponding comparator transition samples an RC-decaying reference to generate an exponential weight, which is then processed by an in-circuit normalization stage. In contrast to analog softmax circuits that rely on transistor weak-inversion behavior for exponentiation, the proposed architecture controls the softmax response through the ramp slope and RC time constant, enabling programmable effective temperature. The 128-element architecture is evaluated using transistor-level and post-layout extracted simulations, including multi-level input vectors, capacitance variation and mismatch, process and temperature variation, monte carlo analysis, and shared-interconnect parasitics. The complete 128-element implementation occupies 9453.42~$μ\mathrm{m}^{2}$ including the shared global ramp circuitry, while each replicated softmax element occupies 70.2~$μ\mathrm{m}^{2}$. The circuit achieves a 242.97-ns evaluation latency at 13.44~mW total power, corresponding to 25.5~pJ per output element. The simultaneous 128-element evaluation achieves an RMSE of 24.46~mV relative to the ideal softmax response. The extracted circuit characteristics are further incorporated into a MemTorch-based hardware-aware Transformer model, where the proposed softmax achieves a validation loss within 2.5\% of the ideal-softmax baseline.

补充信息

↑