VQ-LIC:资源受限FPGA上的共享向量量化学习图像压缩
VQ-LIC: Shared Vector-Quantized Learned Image Compression on a Resource-Constrained FPGA
浏览论文内容
中文总结 AI 辅助
VQ-LIC提出一种在资源受限FPGA上部署学习图像压缩的边云编解码器,通过共享DW/PW引擎实现VQ,以低延迟和高能效达到与大规模编解码器相当的率失真性能。
中文摘要 AI 辅助
学习图像压缩(LIC)难以部署在资源严重受限的FPGA上,因为其实际运行速度不仅取决于算术运算数量,还取决于内存流量、不同操作之间的不平衡以及硬件对工作的批处理方式。我们提出了VQ-LIC,一种非对称的边云编解码器,其中紧凑的INT8深度可分离(DW)-逐点(PW)分析变换和多码本向量量化(VQ)在边缘端运行于可复用的DW/PW引擎对上,而重建则由更大的云端解码器处理。由于VQ码字匹配可表示为点积,因此它直接映射到相同的PW引擎上,无需单独的VQ计算阵列,据我们所知,这是FPGA LIC中的首次。一种新颖的延迟模型,源自FPGA读取、DW、PW和写入成本的确定性RTL周期计数,可预测分析变换的每块延迟;由于VQ共享相同的PW数据路径,该模型也适用于VQ。该模型直接针对硅片验证,预测部署的分析和VQ延迟分别在0.26%和0.05%以内,并指导选择三块$16$-$48$-$64$变换。训练后码本缩减将VQ算术和码本存储减少$4\times$,并缩小固定宽度潜在表示。在220-DSP Zynq-7020上,VQ-LIC的中速率预设达到0.1398比特/像素,PSNR为28.69 dB,MS-SSIM为13.06 dB(CLIC 2017),优于类似规模的神经编码器,并达到比其大三个数量级的编解码器的率失真范围。完整的0.1945-kMAC/像素分析-VQ流水线在硅片上以47.98帧/秒运行,每帧能耗42.84 mJ,使用的DSP数量比同类FPGA LIC加速器少一个数量级,同时以适度的PSNR权衡实现更低的比特率、更高的吞吐量和更低的每帧能耗。
英文摘要
Learned image compression (LIC) is hard to deploy on severely resource-constrained FPGAs, since how fast it actually runs depends not just on arithmetic count, but also on memory traffic, imbalance between different operations, and how the hardware batches its work. We present VQ-LIC, an asymmetric edge-cloud codec in which a compact INT8 depthwise (DW)-pointwise (PW) analysis transform and multi-codebook vector quantization (VQ) run at the edge on a reusable DW/PW engine pair, while reconstruction is handled by a larger cloud decoder. Since VQ codeword matching is expressible as a dot product, it is mapped directly onto the same PW engine, removing the need for a separate VQ compute array, to our knowledge a first for FPGA LIC. A novel latency model, derived from deterministic RTL cycle counts of an FPGA's read, DW, PW, and write costs, predicts an analysis transform's per-block latency; since VQ shares the same PW datapath, the model applies to VQ as well. Validated directly against silicon, the model predicts deployed analysis and VQ latency within 0.26\% and 0.05\%, and guides the selection of a three-block $16$-$48$-$64$ transform. Post-training codebook reduction then cuts VQ arithmetic and codebook storage by $4\times$ and shrinks the fixed-width latent representation. On a 220-DSP Zynq-7020, VQ-LIC's mid-rate preset reaches 0.1398 bits per pixel at 28.69 dB PSNR and 13.06 dB MS-SSIM on CLIC~2017, outperforming a similarly sized neural encoder and reaching a rate-distortion range comparable to a codec three orders of magnitude larger. The complete 0.1945-kMAC/pixel analysis-VQ pipeline runs at 47.98 frames per second and 42.84 mJ per frame on silicon, using an order of magnitude fewer DSPs than comparable FPGA LIC accelerators while achieving lower bitrate, higher throughput, and lower energy per frame at a modest PSNR tradeoff.
发表机构
- School of Science and Engineering (SBASSE), Lahore University of Management Sciences (LUMS)(科学与工程学院,拉合尔管理科学大学)
机构由 AI 辅助整理,请以论文原文为准。