FlowTT:在不规则张量-列车嵌入中利用计算流重用
FlowTT: Exploiting Computation Flow Reuse in Irregular Tensor-Train Embedding
浏览论文内容
中文总结 AI 辅助
FlowTT是一种感知计算流的GPU执行框架,通过优化索引分组、执行路径和调度减少冗余操作与内存开销,在Meta的合成推荐基准上相比EcoRec显著降低了TT嵌入的推理和训练延迟。
中文摘要 AI 辅助
张量-列车(Tensor-Train, TT)分解可有效压缩推荐模型中的大嵌入表,但基于TT的嵌入查找效率仍较低,原因在于输入索引间部分共享的计算流未被完全重用,且中间结果在连续TT核收缩间会被重复从片外存储。我们提出FlowTT,这是一种感知计算流的GPU执行框架,它将TT gather(收集)重新表述为一组前缀共享的不规则计算流。FlowTT结合了与计算流对齐的基于前缀的索引分组、带有片上中间结果保留的融合TT嵌入执行路径,以及基于分块工作窃取和L2 checkpointing(检查点)的持久线程调度,以在倾斜工作负载下维持计算重用。通过结合任务形成、数据缓冲和调度与TT gather的结构,FlowTT减少了冗余TT核操作、全局内存流量和负载不平衡。在Meta的合成推荐基准(Meta-240、Meta-480和Meta-788)上,FlowTT相比现有方法始终实现最低延迟;在批量大小为32768时,与EcoRec相比,它在推理时将延迟降低最多42.2%,在训练时降低最多49.2%,同时还实现了最低的推理峰值内存使用量。这些结果表明,暴露前缀共享的计算是高效TT嵌入执行的关键。
英文摘要
Tensor-Train (TT) decomposition effectively compresses large embedding tables in recommendation models, but TT-based embedding lookup remains inefficient because partially shared computation flows across input indices are not fully reused and intermediate results are repeatedly materialized off-chip between sequential TT-core contractions. We present FlowTT, a flow-aware GPU execution framework that reformulates TT gather as a set of prefix-shared irregular computation flows. FlowTT combines flow-aligned prefix-based index grouping, a fused TT-embedding execution path with on-chip intermediate retention, and persistent-thread scheduling with chunk-based work stealing and L2 checkpointing to preserve reuse under skewed workloads. By co-designing task formation, data buffering, and scheduling with the structure of TT gather, FlowTT reduces redundant TT-core operations, global-memory traffic, and load imbalance. On Meta's synthetic recommendation benchmarks (Meta-240, Meta-480, and Meta-788), FlowTT consistently achieves the lowest latency compared to existing methods. At batch size 32,768, it reduces latency by up to 42.2% in inference and 49.2% in training relative to EcoRec, while also achieving the lowest inference peak memory usage. These results show that exposing prefix-shared computation is key to efficient TT-based embedding execution.
发表机构
- Hanyang University(汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。