GPU-CFR:通过将博弈编译为静态数据流和 CUDA 图重放实现 80 倍加速的反事实遗憾最小化
GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay
浏览论文内容
中文总结 AI 辅助
GPU-CFR 通过将博弈编译为静态数据流和 CUDA 图重放,消除内核启动开销,实现比先前 GPU 实现快 80 倍、比 CPU 实现快 258 倍的加速。
中文摘要 AI 辅助
反事实遗憾最小化(CFR)是少数在 CPU 上运行速度仍快于 GPU 的大型数值工作负载之一。每次迭代都会通过通用树接口以数百万个小规模、相互依赖的 gather 和 scatter 步骤扫描一个包含多达数十亿状态的博弈树。在 GPU 上,每个内核都在微秒级完成,因此内核启动和框架调度主导了运行时间,先前的 GPU 实现输给了优化的 CPU 代码。我们观察到,对于固定博弈,CFR 迭代中除数值外的所有内容在第一次迭代运行前都是已知的。我们提出 GPU-CFR,一个基于此观察构建的编译器和运行时。它将任何博弈一次性编译为静态数据流:扁平边和信息集数组、预计算索引和深度级批量传递固定了整个操作序列,迭代之间仅求解器状态发生变化。静态机会折叠、深度级执行块和双通道可达缓冲区将框架操作数量减少了最多 18.1 倍。由于形状、索引和缓冲区地址从不改变,CUDA 图重放记录一次迭代并以单次图启动重放。在一张 A100 上,跨越包含纸牌游戏、骰子游戏和棋盘游戏的八游戏套件中,GPU-CFR 比同一加速器上最快的先前 GPU CFR 快 29.8 至 80.4 倍,在四个最大的游戏上比最快的开源 CPU 实现之一 LiteEFG 快 14 至 258 倍。编译表示承担了大部分优势:在无加速器的八 CPU 线程上,它已经比 GPU 基线快 2.2 至 51.1 倍。在 CPU 上,优化路径逐位复现参考迭代,树构建和图捕获在首次求解内即可收回成本。GPU-CFR 在不改变更新规则的情况下,在套件的中大型游戏上击败了所有 CPU 和 GPU 基线。
英文摘要
Counterfactual regret minimization (CFR) is one of the few large numerical workloads that still runs faster on CPUs than on GPUs. Each iteration sweeps a game tree with up to billions of states in millions of small, interdependent gather and scatter steps issued through a generic tree interface. On a GPU every kernel finishes in microseconds, so kernel launches and framework dispatch dominate the run time, and prior GPU implementations have lost to optimized CPU code. We observe that for a fixed game, everything about a CFR iteration except the numerical values is known before the first iteration runs. We propose GPU-CFR, a compiler and runtime built on this observation. It compiles any game once into static dataflow: flat edge and information-set arrays, precomputed indices, and depth-level batched passes fix the entire operation sequence, and only solver state changes between iterations. Static chance folding, depth-level execution blocks, and a dual-lane reach buffer cut the number of framework operations by up to 18.1x. Because shapes, indices, and buffer addresses never change, CUDA Graph Replay records the iteration once and replays it with a single graph launch. On one A100, across an eight-game suite that spans card games, dice games, and board games, GPU-CFR runs 29.8--80.4x faster than the fastest prior GPU CFR on the same accelerator, and 14--258x faster than LiteEFG, one of the fastest open-source CPU implementations, on the four largest games. The compiled representation carries most of that margin: on eight CPU threads with no accelerator it is already 2.2--51.1x faster than the GPU baseline. On the CPU the optimized path reproduces the reference iterates bitwise, and tree construction and graph capture pay for themselves within the first solve. GPU-CFR beats every CPU and GPU baseline on the mid-to-large games of the suite without changing the update rule.
发表机构
- Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院)
机构由 AI 辅助整理,请以论文原文为准。