ERASE:用于现代推荐系统更快训练的早期反向传播调度方案
ERASE: EaRly bAckpropagation SchEdule for Faster Training of Modern Recommendation Systems
查看机构详情
- Carnegie Mellon University(卡内基梅隆大学)
- Meta Inc.(元公司)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对轻量级代理模型训练时加速器利用率低的问题,提出ERASE方案,通过分离子图反向传播与前向工作重叠,使大规模点击率模型训练吞吐量提升最高9.51%且保持性能稳定。
中文摘要 AI 辅助
轻量级代理模型可实现快速实验,无需重复训练前沿规模的系统,但其小型内核常导致现代加速器未被充分利用。传统训练将前向传播与反向传播作为不相交阶段调度,加剧了这种低效,使得某一阶段的空闲容量无法被另一阶段的任务填充。我们将前向-前向(FF)算法的分离机制重新解释为一种调度原语:给定局部目标,分离模块的输出会移除下游梯度依赖,使其反向传播在前向传播完成后即可准备就绪。ERASE 在独立 CUDA 流上提前启动每个分离子图的反向传播,将其与后续前向工作重叠。对轻量级 Transformer 的执行轨迹验证了这种重叠及其限制:使设备饱和的内核无并发容量。在大规模点击率模型上,分离六个密集子架构可使训练吞吐量提升高达9.51%,同时归一化熵接近基线。
英文摘要
Lightweight proxy models enable rapid experimentation without repeatedly training frontier-scale systems, but their small kernels often leave modern accelerators underutilized. Conventional training compounds this inefficiency by scheduling the forward and backward passes as disjoint phases, so spare capacity in one cannot be filled by work from the other. We reinterpret the detachment mechanism of Forward-Forward (FF) as a scheduling primitive: given a local objective, detaching a block's output removes downstream gradient dependencies, making its backward pass ready when its forward pass finishes. ERASE launches each detached subgraph's backward pass early on a separate CUDA stream, overlapping it with subsequent forward work. Execution trace on a lightweight transformer demonstrates this overlap and its limit: a kernel that saturates the device leaves no capacity for concurrency. On a large-scale click-through-rate model, detaching six dense subarchitectures improves training throughput by up to $9.51\%$ while keeping normalized entropy close to the baseline.