发表机构
Ampere Computing(安培计算)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种基于加权位堆的延迟感知压缩树生成方法,利用列内加法交换性和Dadda式拓扑优化Radix-4 Booth乘法器及融合点积,相比Wallace树降低了延迟和逻辑体积。
AI 中文摘要
压缩树历来是多种算术单元(尤其是乘法器)的关键组成部分。后者在很大程度上是独立单元,在常见情况下,综合工具可以生成非常接近最优的解决方案。随着大型矩阵乘法结构在计算行业中占据更核心的地位,显然需要更好地理解这些压缩树。此外,其使用边界需要重新定义:点积、非线性算子的多项式拟合,都是乘积之和。由于乘积是部分积之和,最终我们只是将许多东西作为位堆(bit heaps)相加;即压缩树的基本作用。因此,我们必须将它们视为不仅仅是乘法器的组成部分。本文提出了一种基于加权位堆的延迟感知压缩树方法。每个位由其有效性和相对到达时间描述,生成器使用技术无关的归一化延迟单位来估计布尔逻辑压力。通过利用每列内加法的交换性,该方法在缩减阶段之间重新排序位,并使用具有灵活CSA3:2运算符的Dadda式最大带宽拓扑。对于11至64位的Radix-4 Booth乘法器,生成的树显示出比评估的Wallace式最大压缩树更低的建模延迟和估计逻辑体积。融合的四路整数点积进一步展示了相同的表示如何评估微架构约束的成本,例如保留单个乘积冗余格式。这些结果表明,加权位堆生成为探索独立乘法之外的压缩树提供了实用的技术无关基础。
英文摘要
Compression trees have been historically a key component of several arithmetic units, most notably multipliers. The latter have been for the most part standalone units and in common situations synthesis tools can generate very close to optimal solutions. With large matrix multiply structures taking more of a central role in the compute industry, the need for a better understanding of these compression trees is apparent. Furthermore, the boundary of its usage needs to be redefined: dot products, polynomial fitting of non-linear operators, are all sums of products. And since products are sums of partial products, in the end, we are just adding many things together as bit heaps; i.e. the fundamental role of a compression tree. Therefore, we must see them as more than just a component of a multiplier. This paper presents a delay-aware compression-tree methodology based on a weighted bit heap. Each bit is described by its significance and relative arrival time, and the generator uses technology-independent normalized delay units to estimate Boolean logical pressure. By exploiting the commutativity of addition within each column, the method reorders bits across reduction stages and uses a Dadda-style maximum-bandwidth topology with flexible CSA3:2 operators. For Radix-4 Booth multipliers from 11 to 64 bits, the resulting trees show lower modeled delay and estimated logic volume than the evaluated Wallace-style maximum-compression trees. Fused four-way integer dot products further demonstrate how the same representation can evaluate the cost of microarchitectural constraints, such as preserving individual product redundant formats. These results show that weighted bit-heap generation provides a practical technology-independent basis for exploring compression trees beyond standalone multiplication.
Comments8 pages, 12 figures