发表机构
Shopify(Shopify)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TreeWalker 通过偏特化求值对分组树集成推理进行优化,每组遍历一次树,跳过空子树,在多种配置下显著加速,并保持数值精度。
AI 中文摘要
许多推理工作负载会在共享特征值的行组上评估训练好的树集成:离散时间生存模型将每个患者扩展为 $G$ 个时间步,点击率模型对搜索会话中的每个项目进行评分,情景分析在固定其余输入的同时变化少量输入。标准推理将每一行独立处理,并重复共享工作 $G$ 次。我们提出 TreeWalker,将偏特化求值应用于分组推理:常量特征为静态,变化特征为动态。它每组遍历每棵树一次,在变化的分裂点处划分行位掩码,并跳过空子树。训练保持不变:TreeWalker 读取标准的 LightGBM 和 XGBoost 模型。我们证明了一种结构性的工作分解:每棵树的工作分为常量投影子树大小 $|T_c|$、$G$ 次叶子写入以及谓词掩码供给成本 $Q$。对于轨迹评估器,随着 $G \to \infty$,每行工作接近行独立遍历的 $(d_v+1)/(d+1)$ 比例。在 Intel 上,TreeWalker 在参考配置($T=500$, $L=8$)下比行独立遍历快 2.5-3.2 倍,在生存数据集上 $G=128$ 时快 6.8-7.8 倍,在 Arm 上增益更大。在情景分析基准上,它在两种架构的所有 16 种配置中均更快。对于 f64 模型,输出与 treelite 的 GTIL 在求和顺序上一致;对于 f32 模型,f64 累加在 99.98% 的行上比原生 f32 更接近 Kahan 补偿参考,且从未更远。
英文摘要
Many inference workloads evaluate a trained tree ensemble on row groups that share feature values: discrete-time survival models expand each patient into $G$ time steps, click-through-rate models score every item in a search session, and scenario analyses vary a few inputs while holding the rest fixed. Standard inference treats each row independently and repeats the shared work $G$ times. We present TreeWalker, which applies partial evaluation to grouped inference: constant features are static, varying features dynamic. It walks each tree once per group, partitions a row bitmask at varying splits, and skips empty subtrees. Training is unchanged: TreeWalker reads standard LightGBM and XGBoost models. We prove a structural work decomposition: per-tree work splits into the constant-projected subtree size $|T_c|$, $G$ leaf writes, and a predicate-mask provisioning cost $Q$. For the trace evaluator, per-row work approaches a $(d_v+1)/(d+1)$ fraction of a row-independent walk as $G \to \infty$. On Intel, TreeWalker is 2.5-3.2$\times$ faster than a row-independent traversal at the reference configuration ($T=500$, $L=8$) and 6.8-7.8$\times$ faster at $G=128$ on the survival datasets, with larger gains on Arm. On a scenario-analysis benchmark it is faster in all 16 configurations on both architectures. For f64 models, outputs match treelite's GTIL up to summation order; for f32 models, f64 accumulation is closer to a Kahan-compensated reference than native f32 on 99.98% of rows and never farther.
Comments23 pages, 9 figures, 11 tables. Accepted at NeurIPS 2026. Code and data: https://github.com/Shopify/treewalker