arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19859cs.AR

面向融合 Posit 算术的高频多推测乘累加单元

High-frequency Multispeculative Multiply-Accumulation Unit for Fused Posit Arithmetic

Mario Alonso, Miguel Ángel Sacristán, Guillermo Botella, Alberto A. Del Barrio

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种面向32/64位Posit算术的高频多推测乘累加单元,通过重构流水线、采用Booth-4与Kogge-Stone乘法器、以多推测加法器替代宽quire累加器,实现面积最多减少19.8%、能耗降低超50%,周期时间较64位quire方案缩短79.0%。

中文摘要 AI 辅助

Posit 算术为 IEEE 754 浮点标准提供了一种引人注目的替代方案,具有更高的精度。其融合乘累加操作避免了中间舍入,通过 quire(一种覆盖格式全部动态范围的宽定点累加器,以防止长累加过程中的精度损失和溢出)确保精确的数值可重现性。然而,集成如此大的累加器会带来显著的面积和功耗开销。本文提出了一种针对 32 位和 64 位 Posit 的优化、高频多推测 PositMAC 架构。首先,对流水线进行重构以平衡不同阶段。其次,评估了高速乘法拓扑,结果表明采用 Kogge-Stone 加法器的 Booth-4 方案能够满足严格的 0.5ns(2GHz)目标。最后,将宽的整体式 quire 累加器替换为多推测加法器,与基线相比,面积最多减少 19.8%,能耗降低超过 50%。与其他最先进的设计相比,我们的方案实现了最高的运行频率,并且与支持 64 位 quire 的替代方案相比,周期时间最多缩短 79.0%。这一性能的取得并未增加资源开销,因为该设计在面积上严格更小,并且与所有支持 quire 的同类设计相比,每周期能耗更低。

英文摘要

Posit arithmetic offers a compelling alternative to the IEEE 754 floating-point standard, providing enhanced accuracy. Its fused multiply-accumulate operations avoid intermediate rounding, ensuring exact numerical reproducibility through the quire, a wide fixed-point accumulator spanning the format's full dynamic range to prevent precision loss and overflow during long accumulations. However, integrating such large accumulators incurs significant area and power overheads. This paper presents an optimized, high-frequency Multispeculative PositMAC architecture for 32- and 64-bits Posit. First, the pipeline is restructured to balance the different stages. Second, high-speed multiplication topologies are evaluated, showing that a Booth-4 scheme with Kogge-Stone adders meets a stringent 0.5ns target (2Ghz). Finally, the wide monolithic quire accumulator is replaced with a Multispeculative Adder, diminishing area up to 19.8\% while reducing energy consumption by more than 50\% when compared to the baseline. Compared to other state-of-the-art designs, our proposal achieves the highest operating frequency and reduces cycle time by up to 79.0\% with respect to 64-bit quire-enabled alternatives. This performance is attained without increasing resource overhead, as the design remains strictly smaller in area and achieves lower per-cycle energy consumption than all quire-capable counterparts.

发表机构

  • Facultad de Informática, Universidad Complutense de Madrid(马德里康普顿斯大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑