arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨多GPU的单一全局束搜索:十亿记录前沿规模下的高吞吐束搜索

One Global Beam Across Many GPUs: High-Throughput Beam Search at Billion-Record Frontier Scale

Ivan Litvak

arXiv 2610.06718首次发表:更新:

AI 中文总结

本文提出一种跨多GPU的分布式束搜索方法,通过全局归约和竞争无关路由实现十亿规模高吞吐,并在H200和RTX-3060上验证了性能与加速比。

AI 中文摘要

束搜索反复生成大量子节点,去除重复项,并保留最佳的$B$个结果。我们展示了即使保留集合无法容纳于单个设备上,多块GPU也能作为一个整体搜索来执行这些步骤。状态和候选数据保留在GPU上;CPU仅接收小的控制和祖先记录。我们证明,一个抽象的分布式流水线,只要其实现针对每个键执行一次全局归约且路由无竞争,就能返回与单机缩减键top-$B$相同的结果,该结果基于相同的完整候选、整数分数、Hash128等价性以及完全指定的物理布局平局顺序。对历史单T4引脚的静态审计留下了跨缓冲区唯一性问题未解决,并识别出一个独立的多秩散射风险;两者均非复现的故障,其他修订版本在未经比较的情况下既不继承这些缺陷,也不保证正确性。在八块H200 GPU上,一次Cube4运行在$B_{\ m eff}=2{,}900{,}361{,}216$下完成了饱和深度8,耗时$931.266$秒:即$69{,}608{,}669{,}184$个名义父代-生成器对,或推导出的每秒$74.746$百万对。该谜题仍未解决,深度9在用户请求下停止。另一次双T4 Megaminx运行测得在$B_{\ m eff}=82{,}837{,}504$下每秒$30.274$百万个逻辑子节点。这些任务和硬件不同,因此不能确立强扩展性。此外,在一台八RTX-3060主机上,固定计数的八GPU加速比在常见执行配置下为$7.487$,在选定的稳定配置下为$5.605$;弱实际工作吞吐量增益分别为$6.113$和$7.295$。这些对配置敏感的比率使用各自系列的单GPU基线。

英文摘要

Beam search repeatedly makes many children, removes duplicates, and keeps the best $B$. We show how many GPUs can perform these steps as one search even when the retained set does not fit on one device. States and candidates stay on the GPUs; the CPU receives only small control and ancestry records. We prove that an abstract distributed pipeline returns the monolithic reduced-key top-$B$ for the same complete candidates, integer scores, Hash128 equivalence, and a fully specified physical-layout tie order, provided that its implementation performs one global reduction per key and race-free routing. A static audit of the historical one-T4 pin leaves cross-buffer uniqueness unresolved and identifies a separate multi-rank scatter risk; neither is a reproduced failure, and other revisions inherit neither defects nor correctness without comparison. On eight H200 GPUs, a Cube4 run completed saturated depth 8 at $B_{\rm eff}=2{,}900{,}361{,}216$ in $931.266$ s: $69{,}608{,}669{,}184$ nominal parent--generator pairs, or a derived $74.746$ million pairs/s. A separate two-T4 Megaminx run yielded a derived $30.274$ million nominal parent--generator pairs/s at $B_{\rm eff}=82{,}837{,}504$. Those tasks and hardware differ, so they do not establish strong scaling. Separately, on one eight-RTX-3060 host, fixed-count eight-GPU speedup was $7.487$ under a common execution profile and $5.605$ under selected stable profiles; weak actual-work throughput gain was $6.113$ and $7.295$. These profile-sensitive ratios use each series' own one-GPU baseline.

Comments18 pages, 15 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑