arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量化HCLs对固定微架构MXFP4加速器的影响

Quantifying the Effect of HCLs on a Fixed-Microarchitecture MXFP4 Accelerator

Daniele Passaretti, Sajjad Tamimi, Nicola Dall'Ora

arXiv 2609.18792首次发表:更新:

发表机构

Otto-von-Guericke University Magdeburg; Università degli Studi Guglielmo Marconi(奥托·冯·格里克马格德堡大学; 古列尔莫·马可尼大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过固定微架构的MXFP4点积设计,比较多种HCLs与手写RTL,发现HCLs在面积和时序上可匹配或优于手写,选择应基于生态与接口而非QoR。

AI 中文摘要

硬件构造语言(HCLs)旨在提高硬件设计生产力,同时在不改变设计者微架构的情况下生成寄存器传输级(RTL)电路。然而,大多数HCLs之间的比较要么是定性的,要么是在不同设计之间评估结果质量(QoR),这使得难以将语言效应与设计效应分开。本文使用相同的固定设计——OCP MXFP4块点积,这是边缘物理AI推理核心的量化原语,实现为单个12级、II=1流水线,对最广泛使用的HCLs进行了比较。一个SystemVerilog基线之后,是使用Chisel、SpinalHDL、Amaranth、Clash、Bluespec和C++进行高层次综合(HLS)的实现。每个变体在相同的Artix-7器件集上以100 MHz频率通过相同的流程,并由RISC-V软核驱动。在微架构保持不变的情况下,比较是干净的:每个变体都满足时序要求,HCLs在面积上匹配甚至低于手写RTL。剩余的差异并非源于算法,而是源于每个后端如何降低算术运算,以及一个静默切换DSP推断的宽度选择。与HLS(设计决策仅限于pragma)不同,HCLs实现了可比的面积和时序。因此,选择归结为生态系统适配和接口需求,而非QoR。

英文摘要

Hardware Construction Languages (HCLs) aim to improve hardware design productivity while generating register-transfer-level (RTL) circuits without changing the designer's microarchitecture. However, most comparisons between HCLs are either qualitative or evaluate quality of results (QoR) across different designs, making it difficult to separate language effects from design effects. This paper compares the most widely used HCLs using the same fixed design, the OCP MXFP4 block dot product, a quantization primitive at the heart of edge Physical-AI inference, implemented as a single 12-stage, II=1 pipeline. A SystemVerilog baseline is followed by implementations in Chisel, SpinalHDL, Amaranth, Clash, Bluespec, and C++ for high-level synthesis (HLS). Every variant goes through the same flow on the same Artix-7 device set at 100 MhZ, driven by a RISC-V soft core. With the micro-architecture held constant, the comparison is clean: every variant meets timing, and the HCLs match or even undercut hand-written RTL in area. The remaining differences stem not from the algorithm but from how each back end lowers arithmetic, and from a single width choice that silently toggles DSP inference. Unlike HLS, where design decisions are limited to pragmas, the HCLs achieve comparable area and timing. Therefore, the choice comes down to ecosystem fit and interface needs rather than QoR.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑