Omnisolver:面向伊辛自旋玻璃与QUBO求解器的可扩展接口——添加分布式GPU暴力搜索插件
Omnisolver: An extensible interface to Ising spin-glass and QUBO solvers: adding a distributed GPU brute-force plugin
查看机构详情
- Institute of Theoretical and Applied Informatics, Polish Academy of Sciences(波兰科学院理论与应用物理研究所)
- Quantumz.io Sp. z o. o.(Quantumz.io 有限责任公司)
- Institute of Theoretical Physics, Faculty of Fundamental Problems of Technology, Wrocław University of Science and Technology(弗罗茨瓦夫科技大学基础技术学院理论物理研究所)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
Omnisolver新增分布式GPU暴力搜索插件,可在CUDA GPU上精确穷举求解QUBO与伊辛实例,在8个H100 GPU上N=60时求解约3.15天,为经典启发式算法提供真值 oracle。
中文摘要 AI 辅助
本软件更新为Omnisolver添加了omnisolver-bruteforce插件,该插件是用于在支持CUDA的GPU上对QUBO和伊辛实例进行精确穷举搜索的一级插件。本次更新包含三个组件:将[《计算机物理通讯》260卷,2021年,第107728页]中的单GPU暴力搜索内核打包为具有统一Python和CLI接口的一级Omnisolver插件;基于Ray的分布式框架,该框架将搜索拆分为2^k个固定变量子问题,并将其分发至多GPU和主机,同时配备收集并合并部分结果的控制器;通过补偿增量更新、周期性精确能量重锚定以及最佳缓冲区刷新,实现快速float32基态路径的数值稳定化。当N≥40时,稳定化自动激活,公共采样器API保持不变。在密集随机伊辛实例上,我们在8个NVIDIA H100(96GB)GPU上测量了N=60时的穷举求解时间,达到约3.15天,与经验t_{N+1}=2t_N的倍增规则高度吻合,当N≥44时,分布式采样器实现了相对于单个H100的约8倍完整加速。该插件因此成为Simulated Bifurcation Machine等最先进经典启发式算法的实用真值 oracle。
英文摘要
This software update extends Omnisolver with \texttt{omnisolver-bruteforce}, a first-class plugin for exact exhaustive search of QUBO and Ising instances on CUDA-enabled GPUs. The update contributes three components: packaging of the single-GPU brute-force kernel of [Computer Physics Communications 260, 107728, 2021] as a first-class Omnisolver plugin with a uniform Python and CLI interface, a Ray-based distributed framework that splits the search into $2^{k}$ fixed-variable subproblems and dispatches them across multiple GPUs and hosts, with a controller that gathers and merges partial results, and numerical stabilization of the fast \texttt{float32} ground-state path through compensated incremental updates, periodic exact energy re-anchoring, and best-buffer refresh. Stabilization activates automatically for $N \geq 40$ leaving the public sampler API is unchanged. On dense random Ising instances we measure exhaustive solve times up to $N=60$ on $8\times$ NVIDIA H100 (96 GB) GPUs, reaching $\approx\!3.15$ days at $N=60$ in close agreement with the empirical $t_{N+1}=2\,t_N$ doubling rule, with the distributed sampler reaching the full $\approx\!8\times$ speedup over a single H100 from $N\!\gtrsim\!44$ onwards. The plugin thereby serves as a practical ground-truth oracle for state-of-the-art classical heuristics such as the Simulated Bifurcation Machine.