发表机构
Faculty of Informatics and Data Science, University of Regensburg(雷根斯堡大学信息科学与数据科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种带两级分组处理单元的MSDF加速器,通过四种运行时机制减少U-Net推理的计算量,在保持分割精度的同时降低了数字周期与功耗。
AI 中文摘要
用于脑肿瘤分割的U-Net推理需要数十亿次乘累加操作,这推动了能够动态减少计算量的硬件发展,而非仅依赖固定精度或静态模型压缩。最高有效位优先(MSDF)算术在计算过程中会暴露结果的前导数字,从而能在生成完整数值前做出依赖输出的决策。本文提出了一种用于量化U-Net分割的MSDF加速器,该加速器配备支持带符号INT8操作数和流内偏置累加的两级分组处理单元。四种运行时机制直接作用于输出数字流:ReLU层中的精确早期负检测(END)、分割头中的精确仅符号决策、校准的低阶数字跳过以及校准剪枝。其中两种近似机制在精度约束下离线选择,而执行仅需要轻量控制,且不修改存储的权重。在使用nnU-Net针对BraTS训练的残差U-Net上,所提机制将数字周期减少了38.38%,同时在73个保留案例上实现了80.58%的平均Dice分数,而浮点模型的平均Dice分数为81.20%;仅精确机制就将周期减少了18.79%,且不改变量化输出。在45nm工艺下综合后,该处理单元工作频率为500MHz,面积为0.858mm²,经开关活动标注的功耗分析显示,每个192×192的补丁消耗0.726mJ。预计配备共享激活传输的八输出加速器可实现16.6ms的延迟,每个补丁消耗1.67mJ。
英文摘要
U-Net inference for brain-tumor segmentation requires billions of multiply-accumulate operations, motivating hardware that can reduce computation dynamically rather than relying only on fixed precision or static model compression. Most-significant-digit-first (MSDF) arithmetic exposes the leading digits of a result during computation, enabling output-dependent decisions before the full value is generated. This paper presents an MSDF accelerator for quantized U-Net segmentation with a two-stage grouped processing element supporting signed INT8 operands and in-stream bias accumulation. Four runtime mechanisms operate directly on the output digit stream: exact early negative detection (END) in ReLU layers, exact sign-only decision making in the segmentation head, calibrated low-order-digit skipping, and calibrated pruning. The two approximate mechanisms are selected offline under an accuracy constraint, while execution requires only lightweight control and does not modify the stored weights. On a residual U-Net trained with nnU-Net for BraTS, the proposed mechanisms reduce digit cycles by 38.38\% while achieving a mean Dice score of 80.58\% on 73 held-out cases, compared with 81.20\% for the floating-point model; the exact mechanisms alone reduce cycles by 18.79\% without altering the quantized output. Synthesized in 45~nm, the processing element operates at 500~MHz, occupies 0.858~mm$^2$, and consumes 0.726~mJ per $192\times192$ patch under switching-activity-annotated power analysis. A projected eight-output accelerator with shared activation delivery achieves 16.6~ms latency and 1.67~mJ per patch.