ScaleLUT:一种用于实时多尺度超分辨率的全并行可配置查找表加速器
ScaleLUT: A Fully-Parallel Configurable LUT-Based Accelerator for Real-Time Multi-Scale Super-Resolution
浏览论文内容
中文总结 AI 辅助
ScaleLUT提出一种硬件友好的LUT设计框架与全并行可重构加速器,通过YUV域策略、二次幂内核和旋转集成降低存储与计算,在FPGA上实现实时4K多尺度SR,显著减少资源与功耗。
中文摘要 AI 辅助
实时超分辨率(SR)对于边缘设备仍然具有挑战性,因为基于深度学习的方法需要大量的乘加(MAC)运算、资源和功耗。基于查找表(LUT)的SR通过用表查询替代卷积推理来减少计算量,但现有方法仍存在速度有限、存储开销大以及在不同上采样因子下可扩展性差的问题。我们提出了ScaleLUT,一种面向硬件的LUT设计框架和全并行可重构加速器,用于实时多尺度SR。ScaleLUT结合了硬件友好的YUV域策略、二次幂内核和旋转集成,以提高感受野覆盖率,同时降低LUT维度;除法运算被移位操作替代。这些设计相比最先进的基于LUT的SR方法,将内存减少了18.4%。ScaleLUT支持任意输入分辨率和可配置的x2^n上采样因子,采用深度流水线和大规模并行架构。在Xilinx ZCU102 FPGA上实现,它在300 MHz下实现x2上采样时以95.3 FPS的速度进行实时4K SR。与现有SR加速器相比,ScaleLUT使用的LUT减少至少58.6%,触发器减少41.1%,DSP为零,功耗降低42.0%,同时相比最佳基于CPU的SR实现和先前的基于FPGA的SR加速器,分别提供10倍和1.2倍的加速。这些结果证明了联合LUT算法-硬件协同设计对于实用且节能的边缘SR部署的有效性。
英文摘要
Real-time super-resolution (SR) remains challenging for edge devices because deep-learning-based methods require substantial multiply-accumulate (MAC) operations, resources, and power. Lookup-table (LUT)-based SR reduces computation by replacing convolutional inference with table queries, but existing methods still suffer from limited speed, large storage overhead, and poor scalability across upsampling factors. We present ScaleLUT, a hardware-oriented LUT design framework and fully parallel reconfigurable accelerator for real-time multi-scale SR. ScaleLUT combines a hardware-friendly YUV-domain strategy with power-of-two kernels and rotation ensemble to improve receptive-field coverage while reducing LUT dimensionality; division operations are replaced by shifts. These designs reduce memory by 18.4% over state-of-the-art LUT-based SR methods. ScaleLUT supports arbitrary input resolutions and configurable x2^n upsampling factors using a deeply pipelined, massively parallel architecture. Implemented on a Xilinx ZCU102 FPGA, it achieves real-time 4K SR at 95.3 FPS for x2 upscaling at 300 MHz. Compared with existing SR accelerators, ScaleLUT uses at least 58.6% fewer LUTs, 41.1% fewer flip-flops, zero DSPs, and 42.0% lower power, while delivering 10x and 1.2x speedups over the best CPU-based SR implementation and prior FPGA-based SR accelerators, respectively. These results demonstrate the effectiveness of joint LUT algorithm-hardware co-design for practical and energy-efficient edge SR deployment.
发表机构
- The University of Hong Kong (HKU)(香港大学)
- Southern University of Science and Technology(南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。