AI 中文总结
本文提出一种跨多GPU的分布式束搜索方法,通过全局归约和竞争无关路由实现十亿规模高吞吐,并在H200和RTX-3060上验证了性能与加速比。
AI 中文摘要
束搜索反复生成大量子节点,去除重复项,并保留最佳的$B$个结果。我们展示了即使保留集合无法容纳于单个设备上,多块GPU也能作为一个整体搜索来执行这些步骤。状态和候选数据保留在GPU上;CPU仅接收小的控制和祖先记录。我们证明,一个抽象的分布式流水线,只要其实现针对每个键执行一次全局归约且路由无竞争,就能返回与单机缩减键top-$B$相同的结果,该结果基于相同的完整候选、整数分数、Hash128等价性以及完全指定的物理布局平局顺序。对历史单T4引脚的静态审计留下了跨缓冲区唯一性问题未解决,并识别出一个独立的多秩散射风险;两者均非复现的故障,其他修订版本在未经比较的情况下既不继承这些缺陷,也不保证正确性。在八块H200 GPU上,一次Cube4运行在$B_{\ m eff}=2{,}900{,}361{,}216$下完成了饱和深度8,耗时$931.266$秒:即$69{,}608{,}669{,}184$个名义父代-生成器对,或推导出的每秒$74.746$百万对。该谜题仍未解决,深度9在用户请求下停止。另一次双T4 Megaminx运行测得在$B_{\ m eff}=82{,}837{,}504$下每秒$30.274$百万个逻辑子节点。这些任务和硬件不同,因此不能确立强扩展性。此外,在一台八RTX-3060主机上,固定计数的八GPU加速比在常见执行配置下为$7.487$,在选定的稳定配置下为$5.605$;弱实际工作吞吐量增益分别为$6.113$和$7.295$。这些对配置敏感的比率使用各自系列的单GPU基线。
英文摘要
Beam search repeatedly makes many children, removes duplicates, and keeps the best $B$. We show how many GPUs can perform these steps as one search even when the retained set does not fit on one device. States and candidates stay on the GPUs; the CPU receives only small control and ancestry records. We prove that an abstract distributed pipeline returns the monolithic reduced-key top-$B$ for the same complete candidates, integer scores, Hash128 equivalence, and a fully specified physical-layout tie order, provided that its implementation performs one global reduction per key and race-free routing. A static audit of the historical one-T4 pin leaves cross-buffer uniqueness unresolved and identifies a separate multi-rank scatter risk; neither is a reproduced failure, and other revisions inherit neither defects nor correctness without comparison. On eight H200 GPUs, a Cube4 run completed saturated depth 8 at $B_{\rm eff}=2{,}900{,}361{,}216$ in $931.266$ s: $69{,}608{,}669{,}184$ nominal parent--generator pairs, or a derived $74.746$ million pairs/s. A separate two-T4 Megaminx run yielded a derived $30.274$ million nominal parent--generator pairs/s at $B_{\rm eff}=82{,}837{,}504$. Those tasks and hardware differ, so they do not establish strong scaling. Separately, on one eight-RTX-3060 host, fixed-count eight-GPU speedup was $7.487$ under a common execution profile and $5.605$ under selected stable profiles; weak actual-work throughput gain was $6.113$ and $7.295$. These profile-sensitive ratios use each series' own one-GPU baseline.
Comments18 pages, 15 figures