arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23582cs.ARcs.CLcs.LG

Transformer加速器(TFA):一款用于Transformer推理和机器翻译的宏操作INT8硬件芯片

Transformer Accelerator (TFA): A Macro-Op INT8 Hardware Chip for Transformer Inference and Machine Translation

Shashank

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出了支持Transformer推理和多语言机器翻译的TFA INT8硬件加速器,其功能验证达标,端到端速度较22线程CPU提升约20倍,可实现预训练Transformer的比特精确执行。

中文摘要 AI 辅助

我们提出了Transformer加速器(TFA),这是一款可综合、可参数化的INT8内存到内存引擎,用于Transformer推理。一条时分复用数据通路处理提示处理和自回归生成。TFA通过8个512位宏操作描述符实现矩阵乘法、softmax、RMSNorm、逐元素操作以及复制/收集操作。离线编译的程序通过AXI接口获取、验证和分发,支持编码器、解码器以及编码器-解码器模型。该RTL结合了输出 stationary 乘累加阵列、与DMA和计算重叠的乒乓缓冲区、比特精确的倒数平方根和除法单元、键值缓存和嵌入寻址,以及中止安全的零填充写入引擎。UVM环境将输出与比特精确的黄金模型进行字节级比较。在25项测试和34次受限随机运行中,TFA实现了零不匹配、100%功能覆盖率和94.96%代码覆盖率。我们编译了t5-small编码器-解码器流水线用于英语到法语、德语和罗马尼亚语翻译。在10条多语言谚语上,TFA执行了70320个描述符,并匹配了37.9 MB黄金模型输出,零不匹配。INT8输出在5个句子上逐词匹配浮点参考;其余产生有效替代翻译。随机哈达玛重参数化在各层恢复了约11 dB的每张量INT8信噪比。验证配置实现了比22线程CPU约20倍的端到端加速,而更大的设计预计将每个令牌的能耗降低约1000倍。经过RAM推理重编码后,逻辑面积降至2.73 mm²,该设计在SkyWater sky130上完成了符合设计规则的综合与布局布线。TFA展示了使用紧凑硬件和编译器管理的量化对预训练Transformer进行端到端、比特精确的执行。

英文摘要

We present the Transformer Accelerator (TFA), a synthesizable, parameterizable INT8 memory-to-memory engine for transformer inference. One time-multiplexed datapath handles prompt processing and autoregressive generation. TFA implements matrix multiplication, softmax, RMSNorm, elementwise, and copy/gather operations through eight 512-bit macro-op descriptors. Offline-compiled programs are fetched, validated, and dispatched through AXI interfaces, supporting encoder, decoder, and encoder-decoder models. The RTL combines an output-stationary multiply-accumulate array with ping-pong buffers that overlap DMA and compute, bit-exact reciprocal-square-root and divide units, key-value-cache and embedding addressing, and an abort-safe zero-padding write engine. A UVM environment byte-compares outputs against a bit-exact golden model. Across 25 tests and 34 constrained-random runs, TFA achieved zero mismatches, 100% functional coverage, and 94.96% code coverage. We compiled the t5-small encoder-decoder pipeline for English-to-French, German, and Romanian translation. On ten multilingual proverbs, TFA executed 70,320 descriptors and matched 37.9 MB of golden-model output with zero mismatches. INT8 output matched the floating-point reference token-for-token on five sentences; the rest produced valid alternative translations. Randomized-Hadamard reparameterization recovered about 11 dB of per-tensor INT8 signal-to-noise ratio across layers. The verification configuration achieved about 20x end-to-end speedup over a 22-thread CPU, while larger designs are projected to reduce energy per token by about 1000x. After RAM inference recoding, logic area fell to 2.73 mm2, and the design completed design-rule-clean synthesis and place-and-route on SkyWater sky130. TFA demonstrates end-to-end, bit-exact execution of pretrained transformers using compact hardware and compiler-managed quantization.

发表机构

  • Independent Researcher(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

↑