arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06622cs.DCphysics.comp-phquant-ph

GPU发起的离散模拟分岔:低延迟请求与流式稠密耦合

GPU-Initiated Discrete Simulated Bifurcation: Low-Latency Requests and Streaming Dense Couplings

Yaocheng Chen

首次发表
浏览论文内容

中文总结 AI 辅助

提出基于NVIDIA DOCA GPUNetIO的离散模拟分岔架构,通过持久服务与流式求解分别处理并发请求和超内存稠密模型,显著降低延迟并支持大规模矩阵计算。

中文摘要 AI 辅助

基于GPU的优化面临两个通信瓶颈:协调频繁请求以及传递超出设备内存的稠密模型。我们提出了一种离散模拟分岔(dSB)架构,通过NVIDIA DOCA GPUNetIO同时解决这两个问题。对于常驻模型,一个持久服务接收场更新,在单个GPU线程块内执行每次求解,并返回结果。精确整数耦合求和、GPU工作队列和批量传输使接收-求解-回复路径保持在设备上,无需专用的CPU数据路径核心。在与使用相同求解器的基于套接字的服务器比较中,最大的延迟增益出现在并发负载下。当提供的负载从每秒400千次请求增加到800千次请求时,中位往返延迟仅上升6%。在最高测试负载下,中位和第99百分位延迟分别为189微秒和218微秒,而经过调优的持久CPU代理在重复运行中的相应延迟为288微秒和609微秒。对于大于设备内存的模型,流式求解器在GPU上保留动态状态,并跨副本重用传入的耦合块。它以约307 Gb/s的速度评估千万变量稠密二进制矩阵,通过64-MiB数据包缓冲区消耗12.5-TB逻辑矩阵。在植入实例上的基态恢复以及与参考执行的一致性验证了计算。两种模式共同将dSB扩展到并发请求和超出GPU内存的稠密模型。

英文摘要

GPU-based optimization faces two communication bottlenecks: coordinating frequent requests and delivering dense models that exceed device memory. We present a discrete simulated bifurcation (dSB) architecture that addresses both through NVIDIA DOCA GPUNetIO. For resident models, a persistent service receives field updates, executes each solve within one GPU thread block, and returns the result. Exact integer coupling sums, GPU work queues, and batched transmission keep the receive--solve--reply path on the device without a dedicated CPU data-path core. In comparisons with socket-based servers using the same solver, the largest latency gains occur under concurrent load. As the offered load increases from 400 to 800 thousand requests per second, median round-trip latency rises by only 6\%. At the highest tested load, median and 99th-percentile latencies are 189 and 218~$μ$s, compared with 288 and 609~$μ$s for the tuned persistent CPU proxy across repeated runs. For models larger than device memory, a streaming solver retains dynamical state on the GPU and reuses incoming coupling tiles across replicas. It evaluates ten-million-variable dense binary matrices at approximately 307~Gb/s, consuming a 12.5-TB logical matrix through a 64-MiB packet buffer. Ground-state recovery on planted instances and agreement with reference executions verify the computation. Together, the two modes scale dSB to concurrent requests and dense models beyond GPU memory.

↑