发表机构
Hong Kong Baptist University; Hokkaido University; College of Computer Science and Technology, China University of Petroleum (East China); RIKEN Center for Computational Science; A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC); Hong Kong Polytechnic University; National Institute of Advanced Industrial Science and Technology(香港浸会大学; 北海道大学; 中国石油大学(华东)计算机科学与技术学院; 理化学研究所计算科学中心; 新加坡科技研究局高级智能与计算研究所; 香港理工大学; 产业技术综合研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
edgeFBP是一种面向边缘设备的高效外存CT重建框架,采用端到端流水线和混合精度策略加速反投影内核,在Jetson设备上显著提升速度与能效。
AI 中文摘要
计算机断层扫描(CT)是一种重要的三维成像技术,广泛应用于医学诊断和科学研究。然而,由于计算能力、内存容量和能源预算的限制,在边缘设备上执行CT成像具有挑战性。本文提出了一种高效的CT重建框架,称为edgeFBP,专为Nvidia Jetson片上系统(SoC)设备设计。edgeFBP采用端到端流水线设计,在严格的功耗和内存约束下实现高效的外存图像重建。edgeFBP利用混合精度策略,借助半精度张量核心(TCs)加速瓶颈反投影(BP)内核。与广泛使用的RTK库相比,edgeFBP在Jetson Nano上实现了1.83倍的加速,在Jetson AGX上实现了2.56倍的加速。在严格的25瓦功耗预算下,Jetson Nano上的edgeFBP能效比Nvidia DGX A100高出5至48倍,使得在受限边缘设备上实现数据中心规模的成像成为可能。
英文摘要
Computed Tomography (CT) is an essential 3D imaging technology widely used in medical diagnostics and scientific research. However, performing CT imaging on edge devices is challenging due to limitations in computational power, memory capacity, and energy budget. This paper presents an efficient CT reconstruction framework, called edgeFBP, designed for Nvidia Jetson System-on-Chip (SoC) devices. edgeFBP adopts an end-to-end pipeline design for efficient out-of-core image reconstruction under tight power and memory constraints. edgeFBP utilizes a mixed-precision strategy leveraging half-precision Tensor Cores (TCs) to accelerate the bottleneck back-projection (BP) kernel. edgeFBP achieves a 1.83x speedup over the widely used RTK library on Jetson Nano and a 2.56x speedup on Jetson AGX. Under a strict 25-Watt power budget, edgeFBP on Jetson Nano achieves up to 5-48x higher energy efficiency than an Nvidia DGX A100, enabling datacenter-scale imaging on constrained edge devices.