发表机构
University of Bristol; California Institute of Technology; Imperial College London(布里斯托尔大学; 加州理工学院; 伦敦帝国学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对提升决策树的硬件部署难题,提出FQTree细粒度量化算法与QXGB自动硬件生成框架,在JSC等数据集上实现LUT使用量降低26-57%且精度匹配或提升。
AI 中文摘要
提升决策树(BDTs)广泛应用于延迟敏感型应用,但高效的硬件部署仍具挑战性。现有设计常依赖均匀或手动调整的定点格式,可能引入不必要的硬件成本或精度损失。本研究提出用于BDT细粒度量化感知训练的FQTree算法,以及用于自动硬件生成的QXGB框架。FQTree引入面向硬件的叶值量化方案,采用全局量化步长结合逐树移位,实现紧凑的非负整数叶表示,可控制裁剪/剪枝及偏置折叠以降低数据通路成本。本研究在提升过程中应用该量化,使后续树适配已量化集成模型的误差,再通过基于编译器的流程将训练好的模型转换为低延迟硬件实现。在JSC、MNIST和NID数据集上的结果显示,与最先进的基于FPGA的BDT设计相比,本方法减少26-57%的LUT使用量,同时达到或提升精度。
英文摘要
Boosted decision trees (BDTs) are widely used in latency-critical applications, but efficient hardware deployment remains challenging. Existing designs often rely on uniform or manually tuned fixed-point formats, which can introduce unnecessary hardware cost or accuracy loss. This work presents the FQTree algorithm{https://github.com/ecs-bristol/FQTree} for fine-grained quantization-aware training of BDTs, together with the QXGB framework for automatic hardware generation. FQTree introduces a hardware-oriented leaf-value quantization scheme that uses a global quantization step together with a tree-wise shift, enabling compact non-negative integer leaf representations, controlled clipping/pruning, and bias folding to reduce datapath cost. This work further applies this quantization during boosting so that later trees adapt to the errors of the already-quantized ensemble, and then lowers the trained model into low-latency hardware implementations through a compiler-based flow. Results on JSC, MNIST, and NID show that our method reduces LUT usage by 26-57\% compared with the state-of-the-art FPGA-based BDT designs while matching or improving accuracy.
Commentsaccepted by ASAP'26. Code available at https://github.com/ecs-bristol/FQTree